Double-stage wall painting face image restoration method and device and medium
By constructing a two-stage mural face image repair model, using structural reconstruction network and texture repair network, the problems of insufficient overall structural consistency, lack of realism in local texture details, and poor semantic correlation in mural face image repair are solved, and high-quality mural face image repair effect is achieved.
Patent Information
- Application Number
- CN202510215332.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-17
AI Technical Summary
In the repair of mural face images, there are problems such as insufficient global consistency of the overall structure and lack of visual realism in local texture details, and the semantic correlation between the damaged area and the surrounding undamaged part of the face is not strong.
The dual-stage mural face image repair method is adopted to construct a mural face image repair model based on the structural reconstruction network and the texture repair network. The structure reconstruction network is responsible for repairing the overall structure of the face, while the texture repair network is responsible for refining the texture and color transition of the face.
The overall structural consistency of the face image repair results and the authenticity of local texture details are achieved, the coordination of the shape and proportion between the repair area and the undamaged area is enhanced, and the authenticity of the repair results and the consistency of the artistic style is ensured.
Smart Images

Figure CN120163738A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a dual-stage mural face image restoration method, device and medium. Background Art
[0002] As an important medium for carrying history and culture, murals are widely distributed in religious sites, ancient buildings and ruins around the world. They are important materials for studying human history, art and culture. Among them, the human face in the murals is particularly critical. It not only vividly reproduces the image of the characters, clothing characteristics and social outlook of a specific era, but also contains rich artistic expression and cultural symbolic meaning. However, under the dual influence of natural environmental erosion and human destruction, many human faces in the murals have become powdered, peeled off, faded, or even completely disappeared. Natural factors such as moisture, salt stains, and temperature differences can cause cracks and blurred textures in the facial area, while human carving and improper restoration further aggravate the damage to facial details. These damages make it difficult to reproduce the most expressive and emotionally valuable facial images in the murals, greatly weakening the overall artistic and cultural value of the murals.
[0003] Traditional mural protection methods mainly rely on manual restoration, especially the restoration of human faces in murals. This method is completed by professional cultural relics protection personnel in combination with their own experience and auxiliary equipment, but there are significant deficiencies in actual operation. First, manual restoration is highly dependent on personal judgment, and its subjectivity and accuracy are difficult to guarantee. Especially in the restoration of complex textures and details in the face area, it is easy to miss key features or make misjudgments, resulting in distortion or lack of authenticity of the restored face image. Secondly, manual restoration of the face area is time-consuming and costly, which is difficult to meet the increasingly urgent needs of mural protection. Finally, manual restoration is irreversible. Once the restoration effect is poor or errors occur, it may cause irreparable damage to the mural itself, especially for the face area that carries strong emotions and cultural information, the consequences of damage are even more serious. Therefore, in the face of the high requirements for mural face restoration and the limitations of traditional methods, exploring more accurate, efficient and controllable restoration methods has become an urgent need in the field of mural protection.
[0004] With the rapid development of digitalization and artificial intelligence technology, virtual restoration has gradually become an emerging direction for mural protection, especially in the field of facial image restoration. Through digital technology, damaged facial images in murals can be repaired in a non-invasive way, which not only effectively avoids secondary damage to the mural entity, but also restores the detailed features and structural information of the face in a more efficient and accurate way. In addition, the results of virtual restoration can also be used for cultural communication and display, which is convenient for future generations to study and appreciate. Therefore, the restored facial images can not only restore their artistic and cultural value, but also be used for cultural communication and display.
[0005] In recent years, deep learning technology has made remarkable progress in the field of image inpainting, and a series of inpainting methods based on convolutional neural networks and generative adversarial networks have emerged. However, these methods still have many deficiencies in the task of mural face inpainting. When relevant convolutional neural network-based models inpaint mural faces, it is often difficult to simultaneously take into account the coordination of the overall face structure and the realism of local textures. Many models only perform inpainting from local pixel information, lacking effective modeling of the global features of the face, resulting in a lack of consistency in aspects such as the proportion of facial features and the shape of the contour, and it is difficult to restore the overall image of the face. In addition, the semantic association between the damaged area and the undamaged part of the surrounding face is not fully captured, and the inpainted content often has insufficient fusion with the background or other facial areas, prone to structural incoordination or visual artifacts. These problems limit the application of current methods in mural face image inpainting, and more advanced technologies are needed to improve the inpainting effect. Summary of the Invention
[0006] The purpose of this application is to provide a two-stage mural face image inpainting method, device, and medium, which not only solve the problems of insufficient global consistency of the overall face structure and lack of visual realism of local texture details in the process of mural face image inpainting, but also solve the problem of weak semantic association between the damaged area and the undamaged part of the surrounding face.
[0007] To achieve the above purpose, this application provides the following solutions:
[0008] In a first aspect, this application provides a two-stage mural face image inpainting method, including:
[0009] Construct a mural face image inpainting model based on a structure reconstruction network and a texture inpainting network;
[0010] Obtain a mural face image and a mask image;
[0011] Use the relative total variation algorithm to smooth the mural face image and construct a mural face structure image;
[0012] Construct a mural face input image based on the mural face image and the mask image, and construct a mural face input structure image based on the mural face structure image and the mask image;
[0013] Input the mural face input image, the mural face input structure image, and the mask image into the mural face image inpainting model to obtain the inpainted mural face image.
[0014] Optionally, constructing a mural face image inpainting model based on a structure reconstruction network and a texture inpainting network specifically includes:
[0015] Construct a training set and a test set based on the historical mural face input image, the historical mural face input structure image, the historical mask image, and the historically restored mural face image;
[0016] Construct the structure reconstruction network based on the gated mechanism and the spatial frequency domain multi-scale fusion module;
[0017] Construct the texture restoration network based on the structure-guided diffusion model;
[0018] Adopt the meta-learning algorithm to train and test the structure reconstruction network and the texture restoration network in turn based on the training set and the test set. When the trained structure reconstruction network and texture restoration network meet the set requirements, connect the trained structure reconstruction network and the trained texture restoration network as the mural face image restoration model.
[0019] Optionally, constructing the structure reconstruction network based on the gated mechanism and the spatial frequency domain multi-scale fusion module specifically includes:
[0020] Based on the gated mechanism, perform gated convolution calculations on the historical mural face input image, the historical mural face input structure image, and the historical mask image to obtain a first output feature image; the convolution kernel of the first gated convolution is 7×7;
[0021] Perform three downsampling operations on the first output feature image in sequence to obtain a second output feature image; each of the three downsampling operations sequentially performs a second gated convolution operation and a third gated convolution operation, and the convolution kernels of the second gated convolution and the third gated convolution are both 3×3;
[0022] Based on the spatial frequency domain multi-scale fusion module, perform four spatial frequency domain multi-scale fusion operations on the second output feature image in sequence to obtain a third output feature image; each of the four spatial frequency domain multi-scale fusion operations includes a local branch operation and a global branch operation; the third output feature image includes a local feature image and a global feature image;
[0023] Perform three upsampling operations on the third output feature image in sequence to obtain a fourth output feature image, and construct the structure reconstruction network; the upsampling operation includes a fourth gated convolution operation and a nearest neighbor interpolation operation.
[0024] Optionally, based on the gated mechanism, performing gated convolution calculations on the historical mural face input image, the historical mural face input structure image, and the historical mask image to obtain a first output feature image specifically includes:
[0025] Use the first gated convolution to extract features from the historical mural face input image, the historical mural face input structure image, and the historical mask image to obtain an intermediate feature image;
[0026] Use the first ordinary convolution to extract features from the intermediate feature image to obtain a first feature image and a second feature image; the convolution kernel of the first ordinary convolution is 3×3;
[0027] After activating the first feature image and the second feature image respectively using a non-linear activation function, perform element-wise multiplication to obtain the first output feature image; the non-linear activation function includes the Sigmoid function, the tanh function, the ReLU function, the LeakyReLU function, the ELU function, and the Swish function.
[0028] Optionally, based on the spatial frequency domain multi-scale fusion module, perform four spatial frequency domain multi-scale fusion operations on the second output feature image in sequence to obtain a third output feature image, specifically including:
[0029] The local branch operation includes:
[0030] Use the first depthwise separable convolution to extract features from the second output feature image, and activate the extracted features using the non-linear activation function; the convolution kernel of the first depthwise separable convolution is 3×3;
[0031] Use the second depthwise separable convolution to extract features from the activated features, and then add them to the second output feature image to obtain the local feature image; the convolution kernel of the second depthwise separable convolution is 5×5;
[0032] The global branch operation includes:
[0033] Use the second ordinary convolution to extract features from the second output feature image, and activate the extracted features using the non-linear activation function; the convolution kernel of the second ordinary convolution is 1×1;
[0034] Perform a fast Fourier transform on the activated features to obtain a frequency domain feature image;
[0035] Use the second ordinary convolution to extract features from the frequency domain feature image, and activate the extracted features using the non-linear activation function;
[0036] Perform an inverse fast Fourier transform on the activated features to obtain a spatial feature image, and the spatial feature image is the global feature image;
[0037] Through residual connection, add the local feature image and the global feature image to obtain the third output feature image.
[0038] Optionally, the texture repair network is constructed based on a structure-guided diffusion model, specifically including:
[0039] Based on a denoising network, the historical repaired mural face image is fitted according to the fourth output feature image and the mural face input image, and the texture repair network is constructed; the structure-guided diffusion model includes the denoising network.
[0040] Optionally, based on a denoising network, the historical repaired mural face image is fitted according to the fourth output feature image and the historical mural face input image, and the texture repair network is constructed, specifically including:
[0041] Use a third ordinary convolution to extract features from the fourth output feature image and the historical mural face input image to obtain a fifth output feature image; the convolution kernel of the third ordinary convolution is 3×3;
[0042] Perform two residual operations and one downsampling operation on the fifth output feature image in sequence, and after four consecutive times, obtain a sixth output feature image; the residual operation includes a basic convolution operation;
[0043] Perform two residual operations on the sixth output feature image in sequence to obtain a seventh output feature image;
[0044] Perform two residual operations and one upsampling operation on the seventh output feature image in sequence, and after four consecutive times, obtain an eighth output feature image;
[0045] Perform a third ordinary convolution operation on the eighth output feature image to obtain the historical repaired mural face image.
[0046] Optionally, the basic convolution operation includes: group normalization processing, activation by a non-linear activation function, regularization operation, and ordinary convolution operation.
[0047] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the two-stage mural face image repair method described in any one of the above.
[0048] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the two-stage mural face image repair method described in any one of the above is implemented.
[0049] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0050] The present application provides a two-stage mural face image restoration method, device, and medium. By constructing a mural face image restoration model based on a structure reconstruction network and a texture restoration network, the image restored by the structure reconstruction network is used as a condition for the texture restoration network to guide the generation of the mural face texture, ensuring the consistency of the restored mural face texture and the overall structure, and solving the problem of weak semantic correlation between the damaged area and the undamaged part around the face; when constructing the mural face image restoration model by using the structure reconstruction network, not only can local detail features be extracted to obtain a local feature image, enhancing the authenticity and naturalness of facial details, but also the global features of the mural face can be extracted to obtain a global feature image, enhancing the consistency in the overall structure of the face and the proportion of facial features, solving the problem of insufficient global consistency in the overall structure of the mural face image during the restoration process, and ensuring the coordination in shape and proportion between the restored area and the undamaged area; when constructing the mural face image restoration model by using the texture restoration network, the texture details of the mural face are refined and restored, and the facial skin, facial feature details, and color transition can be restored, making the restored mural face image more natural and realistic in texture details, solving the problem of lack of visual authenticity in local texture details when restoring the mural face, and ensuring the authenticity of the restoration result. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is a schematic flowchart of a two-stage mural face image restoration method provided in an embodiment of the present application;
[0053] Figure 2 It is a mural face image, a mask image, and a mural face structure image provided in an embodiment of the present application; Figure 2 In (a) is a schematic diagram of the mural face image; Figure 2 In (b) is a schematic diagram of the mask image; Figure 2 In (c) is a schematic diagram of the mural face structure image;
[0054] Figure 3 It is an overall architecture diagram of a two-stage mural face image restoration method provided in an embodiment of the present application;
[0055] Figure 4 It is a schematic diagram of constructing a structure reconstruction network based on a gating mechanism and a spatial frequency domain multi-scale fusion module and an operation schematic diagram of the spatial frequency domain multi-scale fusion module provided in an embodiment of the present application;Figure 4 In (a) is a schematic diagram of constructing a structure reconstruction network based on a gating mechanism and a spatial frequency domain multi-scale fusion module; Figure 4 In (b) is a schematic diagram of the operation of the spatial frequency domain multi-scale fusion module;
[0056] Figure 5 This is a schematic diagram of a texture repair network and a residual module constructed based on a structure-guided diffusion model provided by an embodiment of the present application; Figure 5 In (a) is a schematic diagram of a texture repair network constructed based on a structure-guided diffusion model; Figure 5 In (b) is a schematic diagram of the residual module;
[0057] Figure 6 This is a display diagram of a historical mural face input structure image and a historically restored mural face image provided by an embodiment of the present application; Figure 6 In (a) is a display diagram of the historical mural face input structure image; Figure 6 In (c) is a display diagram of the historical mural face input structure image; Figure 6 In (e) is a display diagram of the historical mural face input structure image; Figure 6 In (b) is for Figure 6 In (a) is a display diagram of the historically restored mural face image obtained by restoration; Figure 6 In (d) is for Figure 6 In (c) is a display diagram of the historically restored mural face image obtained by restoration; Figure 6 In (f) is for Figure 6 In (e) is a display diagram of the historically restored mural face image obtained by restoration;
[0058] Figure 7 This is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0059] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0060] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0061] The core goal of mural face image restoration is to maintain the consistency of the overall structure and the rationality of the texture. For face images, the consistency of the overall structure requires the restoration model to accurately capture the global semantic information of the face, enabling the restored area to naturally blend with the undamaged area in terms of contour, facial feature proportions, and layout, and avoiding problems such as morphological distortion. The rationality of the texture, on the other hand, emphasizes more on the restoration of local features, especially in the delicate texture, color transition, and texture direction of the face, ensuring that the restoration result not only retains the original artistic style but also reveals real details. Achieving these goals requires capturing the semantic information of the face structure globally while precisely restoring the local texture details, achieving the unity of overall harmony and detailed vividness. This is also the key difficulty and breakthrough direction in virtual restoration technology.
[0062] This application proposes a two-stage mural face restoration method to achieve high-quality restoration effects through a two-stage design. Specifically, this method constructs a mural face structure restoration model based on a structure reconstruction network and a texture restoration network. This model includes a structure reconstruction stage and a texture restoration stage, which are optimized for the problems of global shape modeling and detailed feature generation in face restoration respectively. In the structure reconstruction stage, the mural face structure restoration model focuses on restoring the overall contour and facial feature layout of the face, ensuring the coordination of the restored area with the undamaged area in terms of shape and proportion. In the texture restoration stage, the mural face structure restoration model refines the texture, color transition, and direction features of the face to generate delicate and natural textures, ensuring that the restoration result is realistic and consistent with the mural art style.
[0063] In an exemplary embodiment of this application, a two-stage mural face image restoration method is provided, as Figure 1 shown, including the following steps 101 to step 105. Among them:
[0064] Step 101, construct a mural face image restoration model based on a structure reconstruction network and a texture restoration network.
[0065] Step 102, obtain a mural face image and a mask image.
[0066] Step 103, use the relative total variation algorithm to smooth the mural face image and construct a mural face structure image.
[0067] Step 104, construct a mural face input image based on the mural face image and the mask image, and construct a mural face input structure image based on the mural face structure image and the mask image.
[0068] Step 105, input the mural face input image, the mural face input structure image, and the mask image into the mural face image restoration model to obtain the restored mural face image.
[0069] In the actual process, the above steps 102 to 104 are specifically as follows:
[0070] The obtained mural face image can be expressed as:
[0071]
[0072] where I gt is the mural face image; is the set of real numbers; H is the height of the mural face image; W is the width of the mural face image; H×W is the spatial dimension; C is the number of channels.
[0073] In an exemplary embodiment, the number of channels C of the obtained mural face image I gt is 3.
[0074] The obtained mask image can be expressed as:
[0075]
[0076] where M is the mask image.
[0077] In an exemplary embodiment, the number of channels C of the obtained mask image M is 1.
[0078] The relative total variation (RTV) algorithm is used to smooth the mural face image, and the constructed mural face structure image can be expressed as:
[0079]
[0080] where S gt is the mural face structure image.
[0081] This application uses the RTV algorithm to preprocess the mural face image, removes the texture information in the image, only retains the edge features, and generates the mural face structure image. This processing significantly reduces the interference of texture information on the repair process and provides the mural face structure image as the training target for the subsequent network.
[0082] In an exemplary embodiment, the number of channels C of the constructed mural face structure image S gt is 3.
[0083] Based on the mural face image I gt and the mask image M, the expression for constructing the mural face input image is:
[0084] I in = I gt e(1 - M);
[0085] where Iin is the input image of the mural face; e is the element-wise product.
[0086] Based on the mural face structure image S gt and the mask image M, the expression for constructing the input structure image of the mural face is:
[0087] S in = S gt e(1 - M);
[0088] where S in is the input structure image of the mural face.
[0089] In this application, through a two-stage mural face restoration method, as Figure 3 shown, the relative total variation algorithm is used to smooth the mural face image, so as to extract the mural face structure image. Then, the structure information of the image is restored by combining the mural face image and the mask image through a structure reconstruction network to obtain the restored structure image. Subsequently, in the forward process, noise is gradually added to the images at time 1, time 2,..., time T - 1, and time T to obtain noise images. According to the restored structure image to guide the texture restoration network, the denoising network is used to perform denoising processes on the noise image at time T, the output image at time T - 1,..., the output image at time 2, and the output image at time 1 respectively to further repair the texture of the image, and finally the restored mural face image is obtained. This application effectively solves the problems of insufficient global consistency and loss of local texture details in the restoration of mural face images by using a mural face image restoration model constructed based on a structure reconstruction network and a texture restoration network. The mural face image restoration model adopted in this application effectively avoids the problems of structure distortion and visual artifacts that appear in related methods while retaining the original artistic style of the mural face, and presents a real and natural visual restoration effect of the mural face, fully verifying the effectiveness and superiority of this application in the field of mural face image restoration.
[0090] In an exemplary embodiment of this application, the above step 101 includes the following steps 201 to 204. Among them:
[0091] Step 201, construct a training set and a test set based on the historical mural face input image, the historical mural face input structure image, the historical mask image, and the historically restored mural face image.
[0092] Step 202, construct a structure reconstruction network based on a gated mechanism and a spatial frequency domain multi-scale fusion module.
[0093] Step 203, construct a texture restoration network based on a structure-guided diffusion model.
[0094] Step 204: Using the meta - learning algorithm, train and test the structure reconstruction network and the texture restoration network based on the training set and the test set in sequence. When the trained structure reconstruction network and the texture restoration network meet the set requirements, connect the trained structure reconstruction network and the trained texture restoration network as the mural face image restoration model.
[0095] In the actual process, this application uses historical data to train and test the mural face image restoration model. The historical data includes: historical mural face input images, historical mural face input structure images, historical mask images, and historically restored mural face images.
[0096] In an exemplary embodiment, the above - mentioned step 201 specifically first selects historical mural face images and historical mask images to construct historical mural face input images and historical mural face input structure images, and then uses the historical mural face input images, historical mural face input structure images, historical mask images, and historically restored mural face images as the training set and the test set. Among them, there are 175 historical mural face images. The face parts in the murals are cropped, and then data augmentation techniques are used to perform operations such as rotation, translation, and inversion on the data set to obtain a total of 9552 mural face images with a resolution of 256×256, as shown in part (a) of Figure 2 ; Select historical mask images. The historical mask images are binary matrices. The damaged areas of the historical mask images are represented as 1, and the background areas are represented as 0. Moreover, the mask images include four different sizes, and the sizes of the four types of masks are 10% - 20%, 20% - 30%, 30% - 40%, and 40% - 50% respectively. There are 2000 images of each type of mask, for a total of 8000 mask images, as shown in part (b) of Figure 2 ; The resolution of the historical mural face input structure images is 256×256, as shown in part (c) of Figure 2 Use 2000 historical mural face input images, 2000 historical mural face input structure images, 2000 historical mask images, and 2000 historically restored mural face images as the test set, and use 6000 historical mural face input images, 6000 historical mural face input structure images, 6000 historical mask images, and 6000 historically restored mural face images as the training set.
[0097] In an exemplary embodiment of this application, as shown in Figure 4As shown in part (a) thereof, the structure reconstruction network adopts a U-Net-like architecture. First, it extracts features through gated convolution, then performs three downsampling operations, and then further extracts features through four bottleneck layers. Next, it restores the image features through three upsampling operations, and finally uses gated convolution to restore the structural image of the mural face. The above step 202 includes the following steps 301 to 304. Among them:
[0098] Step 301: Based on the gated mechanism, perform gated convolution calculations on the historical mural face input image, the historical mural face input structural image, and the historical mask image to obtain the first output feature image. The convolution kernel of the first gated convolution is 7×7. The spatial dimension of the first output feature image is H×W, and the number of channels is C.
[0099] Step 302: Perform three downsampling operations on the first output feature image in sequence to obtain the second output feature image. The three downsampling operations all perform the second gated convolution operation and the third gated convolution operation in sequence. The convolution kernel of the second gated convolution is 3×3, and the stride is 2; the convolution kernels of the third gated convolution are all 3×3, and the stride is 1. During the first downsampling process, the spatial dimension after passing through the second gated convolution is The number of channels is 2C, and then the spatial dimension after passing through the third gated convolution is still The number of channels is 2C.
[0100] After one overall downsampling operation, the spatial dimension of the feature is halved, and the number of channels becomes twice the original. A total of three downsampling operations are performed in the structure reconstruction network. When passing through the second downsampling operation, after the change of the image spatial dimension to The number of channels is 4C, the spatial dimension of the second output feature image is The number of channels is 8C.
[0101] Step 303: Based on the spatial frequency domain multi-scale fusion module, perform four spatial frequency domain multi-scale fusion operations on the second output feature image in sequence to obtain the third output feature image. The four spatial frequency domain multi-scale fusion operations all include local branch operations and global branch operations. The third output feature image fuses the local feature image and the global feature image. When the image performs the spatial frequency domain multi-scale fusion operation, it passes through the spatial frequency domain multi-scale fusion module, and its spatial dimension and the number of channels remain unchanged. Therefore, the spatial dimension of the third output feature image is The number of channels is 8C.
[0102] Step 304: Perform three upsampling operations on the third output feature image in sequence, and then perform feature extraction through the first gated convolution to obtain the fourth output feature image, and construct the structure reconstruction network. Among them, the upsampling operation includes the fourth gated convolution operation and the nearest neighbor interpolation operation. The fourth output feature image is the restored structural image.
[0103] In the actual process, when performing the first upsampling operation, first use nearest neighbor interpolation to sample the spatial features of the third output feature image to twice that of the third output feature pixels, and then use three 3×3 gated convolutions to halve the number of channels of the third output feature image. After each upsampling operation, the spatial dimension of the feature becomes twice the original, and the number of channels is halved, that is, the spatial dimension undergoes a transformation from to and then to and finally to the transformation of H×W, and the number of channels changes from 8C, to 4C, then to 2C, and finally to C. Finally, after the fourth gated convolution operation, the number of channels becomes 3. Therefore, the spatial dimension of the restored structural image is H×W, and the number of channels is 3.
[0104] The lightweight structure reconstruction network constructed in this application can effectively reconstruct the overall structure of the mural face, ensuring that the repair results are consistent globally in terms of facial feature proportions, contour shapes, etc., thus achieving a high-quality face repair effect. This application is based on four spatial frequency domain multi-scale fusion modules, and by combining the information extraction capabilities of the spatial domain and the frequency domain, it effectively captures the global semantic relationships and local detail features of the mural face structure image. The ultimate goal of constructing the structure reconstruction network in this application is to generate the fourth output feature image S out , and the expression is
[0105] In an exemplary embodiment of this application, the above step 301 includes the following steps 401 to 403. Among them:
[0106] Step 401, use the first gated convolution to extract features from the historical mural face input image, the historical mural face input structure image, and the historical mask image to obtain an intermediate feature image.
[0107] Step 402, use the first ordinary convolution to extract features from the intermediate feature image to obtain a first feature image and a second feature image.
[0108] Step 403, after activating the first feature image and the second feature image respectively using a non-linear activation function, perform element-wise multiplication to obtain a first output feature image. Among them, the non-linear activation function includes the Sigmoid function, the tanh function, the ReLU function, the LeakyReLU function, the ELU function, and the Swish function.
[0109] In the actual process, the above steps 401 to 403 are represented by the following formula:
[0110] Gating y,x =∑∑W·I;
[0111] Feature y,x = ∑∑W·I;
[0112] O y,x = φ(Feature y,x )eσ(Gating y,x );
[0113] where Gating y,x is the first feature image; Feature y,x is the second feature image; W is the first ordinary convolution; I is the intermediate feature image; y, x are the pixel values of the first and second feature images obtained by using the first ordinary convolution; O y,x represents the first output feature image, φ is the ELU function, and σ is the Sigmod function.
[0114] Feature y,x uses the Sigmod function to map all the values included in Feature y,x to [0, 1] to represent the importance of each local area.
[0115] In an exemplary embodiment of the present application, as shown in part (b) of Figure 4 , the four spatial frequency domain multi-scale fusion operations in step 303 above all include: a local branch operation and a global branch operation. Among them, the first spatial frequency domain multi-scale fusion operation is as follows:
[0116] The local branch operation includes:
[0117] Performing feature extraction on the second output feature image by using a first depthwise separable convolution, and activating the extracted features by using the ReLU function. The convolution kernel of the first depthwise separable convolution is 3×3.
[0118] Performing feature extraction on the activated features by using a second depthwise separable convolution, and then adding the result to the second output feature image to obtain a local feature image. The convolution kernel of the second depthwise separable convolution is 5×5.
[0119] The operation process of the local branch can be represented by the following formula:
[0120] X local = X + DSConv 5×5 (relu(DSConv 3×3 (X)));
[0121] where DSConv 3×3 (·) is the first depthwise separable convolution, X is the second output feature image; relu(·) is the ReLU function; DSConv 5×5(·) is the second depthwise separable convolution, X local is the local feature image.
[0122] The global branch operations include:
[0123] Use the second ordinary convolution to extract features from the second output feature image, and use the LeakyReLU function to activate the extracted features. The convolution kernel of the second ordinary convolution is 1×1.
[0124] Perform a fast Fourier transform on the activated features to obtain the frequency domain feature image. Since the fast Fourier transform has conjugate symmetry, only half of the frequency domain features of the frequency domain feature image need to be convolved. Therefore, convolution operations cannot be directly performed on complex numbers. The real and imaginary parts of the transformed spectrum need to be concatenated using the concatenation function, that is, the second frequency domain feature image, and the concatenation function is the Concat function.
[0125] Use the second ordinary convolution to extract features from the second frequency domain feature image, and use the non-linear activation function to activate the extracted features. The convolution kernel of the second ordinary convolution is 1×1.
[0126] Perform an inverse fast Fourier transform on the activated features, and then perform feature extraction using the second ordinary convolution to obtain the spatial feature image, which is the global feature image.
[0127] The operation process of the global branch can be expressed by the following formula:
[0128] X f = FFT(LeakyReLU(Conv 1×1 (X)));
[0129] X fc = Concat(real(X f ), imag(X f ));
[0130] Y global = Conv 1×1 (FFT -1 (LeakyReLU(Conv 1×1 (X fc )))+X);
[0131] Among them, Conv 1×1 (·) is the second ordinary convolution; LeakyReLU(·) is the LeakyReLU function; X f is the frequency domain feature image; X fc is half of the frequency domain feature image; Concat(·) is the Concat function; FFT is the fast Fourier transform; FFT-1 is the inverse fast Fourier transform; Y global is the global feature image; real(·) is the real part of the frequency-domain feature; imag(·) is the imaginary part of the frequency-domain feature.
[0132] Through the residual connection, the local feature image and the global feature image are added, and the number of channels is adjusted through the second ordinary convolution to obtain the third output feature image. The spatial dimension of the third output feature image is The number of channels is 8C.
[0133] In an exemplary embodiment of the present application, in order to achieve high-quality structure repair results, four different loss functions are used to constrain and guide the training of the network during the construction of the structure reconstruction network, namely the reconstruction loss, the perceptual loss, the style loss, and the adversarial loss.
[0134] The reconstruction loss is defined as the distance between the fourth output feature image S out obtained by the structure reconstruction network and the mural face input structure image S gt , and the formula is defined as follows:
[0135] L1 = ‖S out - S gt ‖;
[0136] where L1 is the distance between the fourth output feature image and the mural face input structure image.
[0137] The specific formula of the perceptual loss is as follows:
[0138]
[0139] where φ i (·) is the activation function of the i-th layer of the VGG19 model; L per is the perceptual loss function.
[0140] The style loss is constrained using the Gram matrix of the feature map, and the implementation formula is as follows:
[0141]
[0142] where G i (·) is the Gram matrix of the feature map at the i-th layer; L sty is the style loss function.
[0143] In order to obtain more realistic results, the adversarial loss is calculated, and the calculation formula is as follows:
[0144]
[0145] where D(·) is the discriminator for the real data Sgt The judgment result; D(G(·)) is the discriminator's judgment on the generated data G(·); L adv is the adversarial loss function; represents the expected value of a random variable, that is, the average value based on all possible outcomes in the sample space.
[0146] In an exemplary embodiment of the present application, the above step 203 specifically includes:
[0147] Based on the denoising network, a texture repair network is constructed by fitting the historical restored mural face image according to the fourth output feature image and the mural face input image. The structure-guided diffusion model therein includes a denoising network.
[0148] In an exemplary embodiment of the present application, as Figure 5 shown in part (a) of, a splicing operation is performed on the fourth output feature image and the historical mural face input image, and then denoising is performed. The denoising network is a U-Net network architecture. This network first extracts features using ordinary convolution, then performs four downsampling operations, then performs two bottleneck layers to further extract features, performs four upsampling operations to restore the noise features, and then uses ordinary convolution to obtain the predicted noise, and finally obtains the historical restored mural face image through skip connection. Based on the denoising network, a texture repair network is constructed by fitting the historical restored mural face image according to the fourth output feature image and the historical mural face input image, specifically including the following steps 501 to step 505.
[0149] Wherein:
[0150] Step 501: Use a third ordinary convolution to extract features from the fourth output feature image and the historical mural face input image to obtain a fifth output feature image. The convolution kernel of the third ordinary convolution is 3×3, and the stride is 1. The spatial dimension of the fifth output feature image is H×W, and the number of channels is C.
[0151] Step 502: Perform two residual operations and one downsampling operation on the fifth output feature image in sequence. After four consecutive times, a sixth output feature image is obtained. Among them, the residual operation includes a basic convolution operation and a noise embedding operation. After each downsampling operation, the spatial dimension of the feature becomes 1 / 2 times the original, and the number of channels is 2 times the original, that is, the spatial dimension goes from to and then to and finally to The transformation, and the number of channels changes from C to 2C, then to 4C, and finally to 8C. Therefore, the spatial dimension of the sixth output feature image is and the number of channels is 8C.
[0152] Step 503: Perform two residual operations on the sixth output feature image in sequence to obtain the seventh output feature image. The spatial dimension of the seventh output feature image is and the number of channels is 8C.
[0153] Step 504: Perform two residual operations and one downsampling operation on the seventh output feature image in sequence. After performing these operations four times continuously, the eighth output feature image is obtained. After each upsampling operation, the spatial dimension of the feature becomes 2 times the original, and the number of channels is halved, that is, the spatial dimension changes from to then to and finally to The number of channels changes from 8C to 4C, then to 2C, and finally to C. Therefore, the spatial dimension of the eighth output feature image is and the number of channels is C.
[0154] Step 505: Obtain the historically restored mural face image by performing the third ordinary convolution operation on the eighth output feature image, as Figure 6 shown, Figure 6 the (a) part in Figure 6 the (c) part in Figure 6 and the (e) part in Figure 6 are all display diagrams of the historically mural face input structure images. The (b) part in Figure 6 is a display diagram of the historically restored mural face image obtained by restoring the historically mural face input structure image in the (a) part in Figure 6 the (d) part in Figure 6 is a display diagram of the historically restored mural face image obtained by restoring the historically mural face input structure image in the (c) part in Figure 6 the (f) part in Figure 6 is a display diagram of the historically restored mural face image obtained by restoring the historically mural face input structure image in the (e) part in
[0155] In an exemplary embodiment of the present application, the number of channels of the historically restored mural face image is 3.
[0156] In an exemplary embodiment of the present application, the residual operation includes a basic convolution operation and a noise embedding operation, as Figure 5As shown in part (b) thereof, the basic convolution operation includes: group normalization processing, activation by a non-linear activation function, regularization operation, and ordinary convolution operation. Specifically, first perform group normalization processing on the fifth output feature image, and then successively perform activation by the Swish function and dropout of the regularization operation on the fifth output feature image after group normalization processing to obtain a regularized feature image. Finally, use the fourth ordinary convolution to extract features from the regularized feature image to obtain a basic convolution image, that is, the fifth output intermediate feature image. Among them, the convolution kernel size of the fourth ordinary convolution is 3×3. The noise embedding operation is to fully connect the noise sequence table to adjust the dimension and channels of the noise sequence, and finally add it to the fifth output intermediate feature extracted by the network. After addition, perform the basic convolution operation again, and then add it to the feature obtained by extracting the features of the fifth output feature image using an ordinary convolution with a convolution kernel size of 1×1 to obtain a feature image after one residual operation. Among them, the spatial dimension of the fifth output feature image is H×W, and the number of channels is C; the spatial dimension of the feature image after one residual operation is H×W, and the number of channels is 2C.
[0157] In this application, the fourth output feature image obtained by using the structure reconstruction network is used to guide the structure diffusion model in the texture repair network to repair the texture information of the mural face. The diffusion model mainly uses the U-Net network as the denoising network. This network fits the real image during the training process, and then uses the sampling algorithm to iteratively sample and generate the repair result. During the sampling generation process, the denoising network needs to be called multiple times. Therefore, the self-attention module in the U-Net network is deleted to reduce the number of network parameters and the computational amount.
[0158] In an exemplary embodiment of this application, based on the denoising network, the actual process of constructing the texture repair network by fitting the historical repaired mural face image according to the fourth output feature image and the historical mural face input image is as follows:
[0159] The spatial dimensions of the fourth output feature image and the historical mural face input image are 256×256, and the number of channels of both is 3 。 First, use a convolution with a stride of 1 and a convolution kernel size of 3×3 to extract the features of the fourth output feature image and the historical mural face input image. After that, the spatial dimension remains unchanged, which is 256×256, but the number of channels becomes 64. Then, use a convolution with a stride of 2, a padding of 1, and a convolution kernel size of 3×3 for downsampling. After each downsampling, the number of channels remains unchanged.
[0160] Before each downsampling, two residual modules are used to perform two residual operations. Each residual module consists of a basic convolution module and a noise embedding module. The basic convolution module performs basic convolution operations. First, group normalization is performed on the fifth output feature image, then through the Swish function and regularization operations, and finally convolution with a stride of 1, padding of 1, and a convolution kernel size of 3×3 is used to extract features. The noise embedding module adjusts the dimension and channels of the noise sequence through a fully connected layer, and finally adds it to the features extracted by the network. After the feature data is input into the residual module, the number of channels of the feature data is first expanded to n times the original by the basic convolution module, where n is specified by the channel doubling sequence. Then, the noise sequence is embedded through the noise embedding module, followed by another basic convolution operation with the number of channels remaining unchanged. Finally, the extracted feature data is added to the original feature data through a residual connection. The channel doubling sequence is set to [1, 2, 4, 8], and the channel doubling sequence controls the change in the number of output channels of the residual module.
[0161] The denoising network performs four downsampling operations in total. Before downsampling, the spatial dimension of the features is 256×256 and the number of channels is 64. Then the features are input into two residual modules and one downsampling module. At this time, the spatial dimension is halved to 128×128, and the number of channels is 64. Then it is input into two residual modules and one downsampling module again. At this time, the spatial dimension is halved to 64×64, and the number of channels is 128. After a total of four downsampling operations in sequence, the final number of channels becomes 512 and the spatial dimension becomes 16×16. After that, the features obtained by downsampling are input into the bottleneck layer. The bottleneck layer consists of two residual modules. After passing through the bottleneck layer, the number of channels and the spatial dimension remain unchanged, still 512 and 16×16. The features extracted by the bottleneck layer use upsampling operations to gradually restore the image, and nearest neighbor interpolation is used for upsampling. Similar to the downsampling operation, two residual network modules are used to adjust the number of channels before upsampling. After the first upsampling, the number of channels is 512 and the spatial dimension is 32×32. After the second upsampling, the number of channels is 256 and the spatial dimension is 64×64. After the third upsampling, the number of channels is 128 and the spatial dimension is 128×128. After the fourth upsampling, the number of channels is 64 and the spatial dimension is 256×256. During the upsampling process, the features obtained by downsampling are fused with the upsampled features through skip connections. After upsampling, the third ordinary convolution operation is used to restore the features to a 3-channel image.
[0162] The forward process of the diffusion model is a Markov chain process. The original image x0 is gradually added noise in the time step T to be transformed into pure Gaussian noise x T : N(0, 1). After using the reparameterization trick, the forward noise addition process can be written as the following formula:
[0163]
[0164] in, x t is the image at time t; α i is the i-th value of the noise sequence list; is α t The cumulative product of ; N(0,1) is the standard normal distribution.
[0165] During the training process, we let the denoising network fit the original image x0. At the same time, in order to speed up the training, the training loss only considers the masked pixels instead of the entire image. The entire training process uses MSE loss as shown below:
[0166]
[0167] Among them, ∈:N(0,1) represents random Gaussian noise; L is the MSE loss function; f θ is the denoising network; s is the face structure image of the mural; is a noise sequence table.
[0168] The texture restoration network uses the mural face structure image restored by the first stage structure reconstruction network as a condition in the sampling generation process, and guides the texture restoration of the mural according to the diffusion model sampling algorithm. The sampling process starts from the pure noise image x T :N(0,1) starts sampling, and removes noise according to multiple iterations of the sampling algorithm to finally obtain a high-quality image. t and the structural image s as the input of the conditional denoising model, the sampling process can be written as:
[0169]
[0170] in, represents the image predicted using the denoising network; α t is the attenuation coefficient at time t; x t-1 is the image at time t-1.
[0171] The original diffusion model requires thousands of iterations to obtain high-quality images. In order to speed up the sampling speed, we use the sampling algorithm proposed in the Denoising Diffusion Implicit Models (DDIM). The specific implementation is to directly obtain the original latent variable x 1:T Construct a subset Where τ is an increasing subsequence of length S derived from [1,...,T]. The final sampling process can be completed in S sampling steps, significantly reducing the number of steps required for sampling. The final sampling process can be expressed as:
[0172]
[0173] where represents random Gaussian noise; is the cumulative product at time τ i -1; is the cumulative product; is the output image at time τ i -1; is the noise standard deviation; ∈ θ is the predicted noise.
[0174] ∈ θ can be calculated by the following formula:
[0175]
[0176] For the definition is as follows:
[0177]
[0178] where η ∈ [0, 1] is an adjustable parameter; is the cumulative product at time τ i When η = 1, the sampling method of the Denoising Diffusion Probabilistic Models (DDPM) is used, and when η = 0, the DDIM sampling method is used. is the cumulative product at time τ
[0179] The resolution of the mural face image is set to 256×256, and training and evaluation are performed on a dataset with a batch size of 8. This application uses the Adaptive Moment Estimation with Weight Decay (AdamW) optimizer with β1 = 0.9 and β2 = 0.5. Here, β1 is the first-order momentum smoothing coefficient, β2 is the second-order momentum smoothing coefficient, and the learning rate is set to 0.0001. The network effect is verified once every 1 iteration (epoch) of the entire dataset, the weight file is stored once every 10 epochs, and a total of 500000 iterations are performed. And tensorboard is used to visually display the loss and images during training.
[0180] The channel multiplier in the texture repair network is set to [1, 2, 4, 8], and the initial number of channels is set to 64. The AdamW optimizer is used in the model, and the learning rate of the optimizer is fixed at 0.0001. During training, the linear noise table is set to [1×10 -6, 0.01], the number of time steps T is set to 2000. Training was carried out for 200 epochs on the dataset. The weight file is saved every 10 epochs of training, and tensorboard is used to visualize the loss and images during the training process.
[0181] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to repair mural face images. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a two-stage mural face image repair method.
[0182] Those skilled in the art can understand that Figure 7 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.
[0183] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0184] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0185] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0186] The databases involved in the various embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0187] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0188] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A two-stage mural face image restoration method, characterized in that: The dual-stage mural face image restoration method comprises: Construct a mural face image restoration model based on the structure reconstruction network and texture restoration network; Obtain mural face image and mask image; Using a relative total variation algorithm to smooth the mural face image to construct a mural face structure image; constructing a mural face input image based on the mural face image and the mask image, and constructing a mural face input structural image based on the mural face structural image and the mask image; The mural face input image, the mural face input structure image and the mask image are input into the mural face image restoration model to obtain a restored mural face image.
2. The dual-stage mural face image restoration method according to claim 1, characterized in that: Based on the structure reconstruction network and texture restoration network, a mural face image restoration model is constructed, which includes: Construct training sets and test sets based on historical mural face input images, historical mural face input structure images, historical mask images, and historical restored mural face images; Constructing the structure reconstruction network based on the gating mechanism and the spatial frequency domain multi-scale fusion module; Constructing the texture restoration network based on the structure-guided diffusion model; A meta-learning algorithm is used to train and test the structure reconstruction network and the texture restoration network in sequence based on the training set and the test set until the trained structure reconstruction network and the texture restoration network meet the set requirements. The trained structure reconstruction network and the trained texture restoration network are connected as a mural face image restoration model.
3. The dual-stage mural face image restoration method according to claim 2, characterized in that: The structure reconstruction network is constructed based on the gating mechanism and the spatial frequency domain multi-scale fusion module, specifically including: Based on the gating mechanism, gated convolution calculation is performed on the historical mural face input image, the historical mural face input structure image and the historical mask image to obtain a first output feature image; the convolution kernel of the first gated convolution is 7×7; Performing three downsampling operations on the first output feature image in sequence to obtain a second output feature image; the three downsampling operations are all performed in sequence with a second gated convolution operation and a third gated convolution operation, and the convolution kernels of the second gated convolution and the third gated convolution are both 3×3; Based on the spatial frequency domain multi-scale fusion module, the second output feature image is sequentially subjected to four spatial frequency domain multi-scale fusion operations to obtain a third output feature image; the four spatial frequency domain multi-scale fusion operations each include a local branch operation and a global branch operation; the third output feature image includes a local feature image and a global feature image; The third output feature image is sequentially subjected to three upsampling operations to obtain a fourth output feature image, and the structure reconstruction network is constructed; the upsampling operation includes a fourth gated convolution operation and a nearest neighbor difference operation.
4. The dual-stage mural face image restoration method according to claim 3 is characterized in that: Based on the gating mechanism, gated convolution calculation is performed on the historical mural face input image, the historical mural face input structure image and the historical mask image to obtain a first output feature image, specifically including: Using a first gated convolution to perform feature extraction on the historical mural face input image, the historical mural face input structure image, and the historical mask image to obtain an intermediate feature image; Using a first common convolution to extract features from the intermediate feature image to obtain a first feature image and a second feature image; the convolution kernel of the first common convolution is 3×3; After activating the first feature image and the second feature image using a nonlinear activation function, element-by-element multiplication is performed to obtain the first output feature image; the nonlinear activation function includes a Sigmoid function, a tanh function, a ReLU function, a LeakyReLU function, an ELU function and a Swish function.
5. The dual-stage mural face image restoration method according to claim 4, characterized in that: Based on the spatial frequency domain multi-scale fusion module, the second output feature image is sequentially subjected to four spatial frequency domain multi-scale fusion operations to obtain a third output feature image, specifically comprising: The local branch operation includes: Using a first depthwise separable convolution to extract features from the second output feature image, and using the nonlinear activation function to activate the extracted features; the convolution kernel of the first depthwise separable convolution is 3×3; A second depthwise separable convolution is used to extract features from the activated features, and then the features are added to the second output feature image to obtain the local feature image; the convolution kernel of the second depthwise separable convolution is 5×5; The global branch operation includes: Using a second common convolution to extract features from the second output feature image, and using the nonlinear activation function to activate the extracted features; the convolution kernel of the second common convolution is 1×1; Perform fast Fourier transform on the activated features to obtain frequency domain feature images; Using a second ordinary convolution to extract features from the frequency domain feature image, and using the nonlinear activation function to activate the extracted features; Performing a fast inverse Fourier transform on the activated features to obtain a spatial feature image, which is the global feature image; The local feature image and the global feature image are added together through a residual connection to obtain the third output feature image.
6. The dual-stage mural face image restoration method according to claim 3 is characterized in that: The texture restoration network is constructed based on the structure-guided diffusion model, specifically including: Based on the denoising network, the historically restored mural face image is fitted according to the fourth output feature image and the mural face input image to construct the texture restoration network; the structure-guided diffusion model includes the denoising network.
7. The dual-stage mural face image restoration method according to claim 6, characterized in that: Based on the denoising network, fitting the historical restored mural face image according to the fourth output feature image and the historical mural face input image to construct the texture restoration network specifically includes: Using a third common convolution to extract features from the fourth output feature image and the historical mural face input image to obtain a fifth output feature image; the convolution kernel of the third common convolution is 3×3; Performing two residual operations and one downsampling operation on the fifth output feature image four times in sequence to obtain a sixth output feature image; the residual operation includes a basic convolution operation; performing the residual operation twice on the sixth output feature image in sequence to obtain a seventh output feature image; Performing two residual operations and one upsampling operation on the seventh output feature image four times in sequence to obtain an eighth output feature image; By performing the third common convolution operation on the eighth output feature image, the historically restored mural face image is obtained.
8. The dual-stage mural face image restoration method according to claim 7, characterized in that: The basic convolution operation includes: group normalization processing, nonlinear activation function activation, regularization operation and ordinary convolution operation.
9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the two-stage mural face image restoration method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the two-stage mural face image restoration method described in any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Intelligent detection system and device for wheel set tread damage of heavy-load train
CN122023401A