Damaged Qinxi character imitation image generation method and image restoration method
By generating adversarial networks to generate damaged masks to process intact Qin bamboo slips text images, the problem of large differences between counterfeit training samples and real ancient text images in existing technologies is solved, and high simulation and sample expansion of ancient text image restoration are achieved.
Patent Information
- Application Number
- CN202510761491.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-09
AI Technical Summary
When existing technologies use machine learning to repair ancient text images, the imitated training samples are quite different from the real damaged ancient text images, resulting in poor model repair effects.
A generative adversarial network is used to generate damage masks, and all types of image data from the same period are used as reference objects for extracting damage masks. The damage masks are then processed using a generative adversarial network to generate highly realistic images of damaged Qin bamboo slips.
It effectively expands the number of damaged ancient character samples and improves the restoration effect and simulation of ancient character image restoration using machine learning.
Smart Images

Figure CN120672888A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ancient character restoration, and in particular to a method for generating an imitation image of damaged Qin bamboo slip characters and a method for restoring an image of damaged Qin bamboo slip characters. Background Art
[0002] With the development of science and technology, many methods for image restoration of ancient text images have emerged using machine learning methods such as neural networks and generative adversarial networks. Image restoration of ancient text images using machine learning requires a large amount of training data. To obtain a large amount of training data, training samples can be generated through imitation.
[0003] Publication number CN116681604A discloses methods for generating damaged ancient Chinese characters, including image masking, image rotation, image cropping, image scaling, salt and pepper noise, and Gaussian noise. In practice, these methods are essentially standard processing methods. The damaged ancient Chinese characters images generated using these methods differ significantly from authentic damaged ancient Chinese characters. If these simulated training samples are used for machine learning, the resulting model characteristics will not be able to effectively repair authentic damaged ancient Chinese characters.
[0004] It can be seen from this that a more reasonable method of imitating damaged ancient characters is currently needed to meet the requirements of machine learning in related fields. Summary of the Invention
[0005] The present invention provides a method for generating an image of damaged Qin bamboo slips and an image restoration method. The method for generating an image of damaged Qin bamboo slips uses all types of image data from the same period as reference objects for extracting damage masks, extracts damage masks by generating an adversarial damage network, and then processes the intact Qin bamboo slip image with the extracted damage mask, thereby obtaining a highly realistic image of damaged Qin bamboo slip text. This method provides a method that can effectively expand the number of damaged ancient text samples and has good practicality in machine learning applications in related fields.
[0006] Accordingly, the present invention provides a method for generating an imitation image of damaged Qin bamboo slips, comprising:
[0007] Constructing a damaged image dataset, the damaged image dataset including a plurality of damaged training data, each damaged training data including a damaged training image and a corresponding intact training image, the intact training images including intact training images of Qin bamboo slips characters and intact training images of non-Qin bamboo slips characters, and the damaged training images including damaged training images of Qin bamboo slips characters and damaged training images of non-Qin bamboo slips characters;
[0008] Initializing a generative adversarial damage network, which includes a mask generator, a damage generator, and a damage discriminator. The mask generator is configured to output a predicted damage mask based on an input intact image. The damage generator is configured to process the corresponding intact image according to the predicted damage mask to obtain a predicted damaged image. The damage discriminator is configured to evaluate the authenticity of the input damaged image according to a preset damage evaluation index.
[0009] Training a generative adversarial network, iteratively training the generative adversarial network using the damaged training data in the damaged image dataset until the training is completed;
[0010] Constructing a mask dataset, sequentially importing intact training images from the damaged image dataset into a trained generative adversarial network, causing a mask generator in the generative adversarial network to sequentially generate corresponding predicted damage masks, and storing all predicted damage masks generated by the mask generator in the mask dataset;
[0011] The imitation image of damaged Qin bamboo slips is generated by randomly extracting a predicted damage mask from the mask data set to process any intact training image of Qin bamboo slips in the damaged image data set to obtain the corresponding imitation damaged training image, which is the required imitation image of damaged Qin bamboo slips.
[0012] In an optional implementation manner, the damaged training image has a preset fixed resolution.
[0013] In an optional implementation manner, both the damaged training image and the intact training image are binarized images.
[0014] Accordingly, the present invention provides a method for repairing damaged Qin bamboo slips, including a training process and a repair process. The training process includes:
[0015] Constructing a damaged Qin bamboo slips character real image dataset, wherein the damaged Qin bamboo slips character real image dataset includes a plurality of damaged Qin bamboo slips character real image data, each of which includes a damaged Qin bamboo slips character real image training image and a corresponding intact training image;
[0016] Constructing a damaged Qin bamboo slips imitation image dataset, wherein the damaged Qin bamboo slips imitation image dataset includes a plurality of damaged Qin bamboo slips imitation image data, each of which includes a damaged Qin bamboo slips imitation image generated based on the damaged Qin bamboo slips imitation image generation method and a corresponding intact training image;
[0017] Constructing a damaged Qin bamboo slip character image dataset, merging the damaged Qin bamboo slip character real image dataset and the damaged Qin bamboo slip character imitation image dataset to obtain the damaged Qin bamboo slip character image dataset, wherein the damaged Qin bamboo slip character image dataset includes a plurality of damaged Qin bamboo slip character image data, each damaged Qin bamboo slip character image data includes a damaged Qin bamboo slip character training image and a corresponding intact training image, and the types of the damaged Qin bamboo slip character training images include the damaged Qin bamboo slip character real training images and the damaged Qin bamboo slip character imitation training images;
[0018] Initializing a generative adversarial restoration network, which includes a restoration generator and a restoration discriminator. The restoration generator is used to output a predicted intact Qin bamboo slip text image based on an input damaged Qin bamboo slip text image, and the restoration discriminator is used to evaluate the authenticity of the intact Qin bamboo slip text image input to the restoration discriminator.
[0019] Train a generative adversarial repair network, and iteratively train the generative adversarial repair network using the training data in the damaged Qin bamboo slips text dataset until the training is completed.
[0020] In an optional implementation manner, the repair process includes:
[0021] An unknown damaged Qin bamboo slip text is input into the generative adversarial repair network, and the predicted intact Qin bamboo slip text image generated by the generator of the generative adversarial repair network is the repaired Qin bamboo slip text image after the unknown damaged Qin bamboo slip text has been repaired.
[0022] In an optional embodiment, the restoration generator includes a first residual block, a second residual block, a bottleneck convolution layer, a first upsampling module, a second upsampling module, a third upsampling module, a bilinear interpolation layer and a normal convolution layer connected in sequence.
[0023] In an optional implementation manner, the first upsampling module and the second upsampling module are respectively provided with a corresponding multi-scale linear attention module.
[0024] In an optional embodiment, the loss function of the restoration generator includes reconstruction loss, adversarial loss and perceptual loss.
[0025] In an optional implementation manner, the restoration discriminator is a DCGAN discriminator.
[0026] In an optional implementation manner, the loss function of the restoration discriminator includes the binary classification loss of the predicted intact Qin bamboo slip text image and the binary classification loss of the intact training image.
[0027] The present invention provides a method for generating an imitation image of damaged Qin bamboo slips and a method for repairing an image of damaged Qin bamboo slips. The method for generating an imitation image of damaged Qin bamboo slips uses all types of image data from the same period as reference objects for extracting damage masks, extracts damage masks by generating an adversarial damage network, and then processes the intact Qin bamboo slip image with the extracted damage mask, thereby obtaining a highly realistic image of damaged Qin bamboo slips. This method provides a method that can effectively expand the number of damaged ancient character samples and has good practicality in machine learning applications in related fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is a flow chart of the method for generating an imitation image of damaged Qin bamboo slips according to the first embodiment of the present invention.
[0029] Figure 2 This is an example diagram of a damaged training data in Example 1 of the present invention
[0030] Figure 3 This is a flow chart of a method for repairing damaged Qin bamboo slips text images according to an embodiment of the present invention.
[0031] Figure 4 Schematic diagram of the repair generator structure of embodiment 3 of the present invention
[0032] Figure 5 2 is a structural diagram of the residual block ResBlock according to the third embodiment of the present invention.
[0033] Figure 6 2 is a schematic structural diagram of an up-sampling module Upsample according to a third embodiment of the present invention.
[0034] Figure 7 Schematic diagram of the structure of the multi-scale linear attention module MSLA of embodiment 3 of the present invention.
[0035] Figure 8 Schematic diagram of the repair discriminator structure of embodiment 4 of the present invention.
[0036] Figure 9 Schematic diagram of the downsample layer structure of the fourth embodiment of the present invention.
[0037] Figure 10 This is a schematic diagram comparing the physical image of Qin bamboo slips characters and the binary image of Qin bamboo slips characters in Example 5 of the present invention.
[0038] Figure 11 Schematic diagram of a repair example of a repair generator according to the fifth embodiment of the present invention DETAILED DESCRIPTION
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0040] Embodiment 1
[0041] Figure 1 It is a flowchart of a method for generating a forged image of damaged Qin bamboo slips characters according to an embodiment of the present invention.
[0042] An embodiment of the present invention discloses a method for generating a forged image of damaged Qin bamboo slips characters, including
[0043] S101: Construct a damaged image data set;
[0044] The damaged image data set includes a number of damaged training data. Each damaged training data includes a damaged training image and a corresponding intact training image. The damaged training image and the intact training image have the same resolution. The damaged training image includes a damaged training image of Qin bamboo slips characters and a damaged training image of non-Qin bamboo slips characters. The intact training image includes an intact training image of Qin bamboo slips characters and an intact training image of non-Qin bamboo slips characters;
[0045] Figure 2 It is an example diagram of a damaged training data according to an embodiment of the present invention. Specifically, the meaning of this damaged training data is the character 'ke'. The intact training image of Qin bamboo slips characters is a clear binary image of the character 'ke', and the damaged training image is a binary image of the character 'ke' with defects.
[0046] It should be noted that in this step, the damaged training image includes a damaged training image of Qin bamboo slips characters and a damaged training image of non-Qin bamboo slips characters, and the intact training image includes an intact training image of Qin bamboo slips characters and an intact training image of non-Qin bamboo slips characters. That is, although the embodiment of the present invention needs to generate a forged image of damaged Qin bamboo slips characters, when collecting damaged training data, images of other works of the same period can be collected in the same data form. The internal logic of this implementation means is that the image of damaged Qin bamboo slips characters is essentially an image, and the damaged traces are substantially the same as those of non-Qin bamboo slips characters. The purpose of the embodiment of the present invention is to extract the damaged mask, and the damaged mask reflects the damaged characteristics rather than the characteristics of the characters. Therefore, in the stage of collecting training data, the damaged training image includes a damaged training image of Qin bamboo slips characters and a damaged training image of non-Qin bamboo slips characters, and the intact training image includes an intact training image of Qin bamboo slips characters and an intact training image of non-Qin bamboo slips characters.
[0047] S102: Initialize the generated adversarial network;
[0048] The generative adversarial damage network includes a mask generator, a damage generator and a damage discriminator.
[0049] The mask generator is used to output a predicted damage mask based on the input intact image. Specifically, assuming that the input intact training image is I good After processing by the mask generator, a predicted damage mask M is output. It should be noted that since the data format used in this embodiment of the present invention is a binary image, which does not contain fuzzy content, the resulting damage mask is actually an obscured mask. This determines the definition of the damage generator formula. If the data format is grayscale or color, a fuzzy mask is more likely. When using a fuzzy mask, the damage generator formula needs to supplement it with a parameter regarding the strength of the masking effect.
[0050] The damage generator is used to process the intact training image according to the predicted damage mask to obtain a predicted damaged image. Specifically, the predicted damaged image , where the symbol ⊙ represents the Hadamard product, Understood as random masking noise, Indicates that the mask Directly act on the intact image to obtain the first intermediate image after masking; superimpose on the first intermediate image , is the noise Random masking is performed in the areas not masked by the mask, where It can be understood as a compensation parameter that can be adjusted during the training process, so that the model has a certain degree of flexibility during the training process to avoid the situation where the input results and output results are completely equivalent;
[0051] The damage discriminator is used to evaluate the authenticity of the input predicted damage image or the damaged training image according to a preset damage evaluation index.
[0052] Specifically, the damage generator can adopt a U-Net structure, and the damage discriminator can adopt a PatchGAN structure.
[0053] Specifically, the damage generator adopts a U-Net structure, which specifically includes a damage encoder and a damage decoder. Specifically, the damage encoder mainly performs the downsampling path, consisting of multiple convolutional layers, using the convolutional layers to gradually extract high-level semantic features and gradually reduce the spatial resolution of the input intact image. The damage decoder mainly performs the upsampling path, consisting of multiple transposed convolutional layers (for upsampling), using the transposed convolutional layers to gradually restore the spatial resolution of the image and generate the required mask. Specifically, the number of convolutional layers used for downsampling in the damage encoder and the transposed convolutional layers used for upsampling in the damage decoder are strictly corresponding, and skip links need to be set between the damage encoder and the damage decoder. Specifically, the feature map generated by a convolutional layer used for upsampling in the encoder is jump-connected to the device convolution layer of the corresponding level (same resolution) in the decoder to achieve the transmission of low-level detail information (details lost by the encoder during compression are directly transferred to the decoder via skip connections), alleviate gradient vanishing (providing a shortcut for gradients to flow directly from deep layers of the decoder to shallow layers of the encoder), and fuse multi-scale information (features of different scales are fused to make the generated image have global structural consistency).
[0054] Specifically, the damage generator is divided into layers according to the number of downsampling times, and generally 4 to 8 layers are required. Since the intact image of the embodiment of the present invention has low complexity and low resolution (generally 128X128), the layers are generally set to 4 to 5 layers; in each layer, a normal convolution with a 3x3 kernel or a 4x4 kernel and a step size of 1 is set for feature extraction, and then a LeakyReLU or ReLU is set for feature activation, and then downsampling is performed through pooling (parameters are set to 2x2 pooling, step size is 2) or strided convolution (parameters are set to 4x4 convolution, step size is 2); after the downsampling layer of the last layer is completed, 2 to 3 bottleneck convolution layers are generally set for the most abstract feature extraction.
[0055] Specifically, the damage decoder is divided into layers according to the number of upsampling times, and the number of layers is set to be the same as that of the damage generator. In each layer, a transposed convolution (parameters are set to 4x4 kernel, step size is 2) is generally used for upsampling; specifically, the jump link point between the damage encoder and the damage decoder can be set after the transposed convolution of any layer (the feature map resolution of the corresponding link of the damage encoder must be consistent with the resolution of the transposed convolution output). Generally, a jump link is set after the transposed convolution of the last layer; specifically, the fusion method at the jump link adopts the splicing method (optionally the addition method); before the mask is output, after ordinary convolution processing and activation function activation, the required mask is generated by the normalization function (generally Tanh); the purpose of normalization processing is that the mask itself has normalized features (black and white features), and the format of the output data can be well coordinated through normalization processing.
[0056] Specifically, the damage discriminator adopts the PatchGAN structure. PatchGAN is a discriminator architecture designed specifically for image generation tasks (especially image-to-image translation). Its core idea is not to directly judge the authenticity of the entire image, but to independently judge the local area (patch, receptive field) of the image. Finally, it outputs a statistical graph or statistical matrix about the "local authenticity" of each receptive field. Each element in the statistical graph or statistical matrix represents the probability that the corresponding local area (patch) in the input image is a real image.
[0057] Specifically, the receptive field of view of the damage discriminator (i.e., the size of the input image area "seen" by each output neuron) can be selected from 1x1, 16x16, 30x30, etc. The output of the damage discriminator is a statistical matrix, but during training, the average value of all output units is usually taken as the final discriminant loss for the entire image.
[0058] Specifically, the input of the damage discriminator is the concatenation of the predicted damage image and the damaged training image. The damage discriminator will calculate the similarity (true probability) of each receptive field of view based on the receptive field of view. The output result is a statistical matrix, and the average value of the statistical matrix is used as the damage evaluation indicator. The loss function can be a superposition of adversarial loss (cGAN loss) and L1 regularization loss.
[0059] S103: training a generative adversarial network;
[0060] Iteratively training the generative adversarial network using the damaged training data in the damaged image dataset until the training is completed;
[0061] Specifically, the generative adversarial network of the embodiment of the present invention has no essential difference in form from the generative adversarial network of the prior art. The innovative content of the generative adversarial network of the embodiment of the present invention mainly lies in the structural design of the generative adversarial network (with the addition of a mask generator) and the selection of damaged training data (the damaged training images include damaged training images of Qin bamboo slips and damaged training images of non-Qin bamboo slips); therefore, the structures of the damaged generator and the damaged discriminator can be designed based on the existing technology, and the training process can also be carried out according to the existing technology, and the embodiment of the present invention does not impose any additional restrictions.
[0062] S104: constructing a mask dataset;
[0063] Importing intact training images of the damaged training data in the damaged image dataset into the trained generative adversarial damage network in sequence, and using a mask generator in the generative adversarial damage network to generate corresponding predicted damage masks in sequence, and storing the predicted damage masks in a mask dataset;
[0064] After the training of the generative adversarial network is completed, when a good training image is input into the generative adversarial network, the mask generator will derive the corresponding mask; traversing the input good training image, the mask generator will derive a corresponding number of masks accordingly; and the masks derived by the mask generator are stored in a mask dataset.
[0065] S105: Generate a replica image of damaged Qin bamboo slips text, randomly extract a predicted damage mask from the mask data set, and process any intact training image of Qin bamboo slips text in the damaged image data set to obtain a corresponding replica damaged training image, wherein the replica damaged training image is the required replica image of damaged Qin bamboo slips text.
[0066] The masks in the mask data set represent the current damage changes of the physical carriers of the same period. Therefore, the damage changes are actually not corresponding, that is, although the mask is extracted from a corresponding intact training image, in fact any physical carrier of the same period may have corresponding damage conditions; therefore, in this step, different intact training images of Qin bamboo slips are traversed and processed using different masks, and a large number of imitation images of damaged Qin bamboo slips can be imitated. The imitation images of damaged Qin bamboo slips generated in this way have a high degree of simulation, and the model trained by the imitation images of damaged Qin bamboo slips can be more in line with the actual situation in application.
[0067] Example 2
[0068] Figure 3 This is a flow chart of a method for repairing damaged Qin bamboo slips text images according to an embodiment of the present invention.
[0069] An embodiment of the present invention provides a method for repairing a damaged Qin bamboo slip text image, including a training process and a repair process.
[0070] The training process includes steps S201 to S205.
[0071] S201: Construct a dataset of damaged Qin bamboo slips. The dataset includes a plurality of pieces of damaged Qin bamboo slips image data, each piece of damaged Qin bamboo slips image data including a real training image of damaged Qin bamboo slips and a corresponding intact training image.
[0072] S202: Constructing a damaged Qin bamboo slip imitation image dataset. The damaged Qin bamboo slip imitation image dataset includes a plurality of damaged Qin bamboo slip imitation image data, each of which includes a damaged Qin bamboo slip imitation image generated based on the damaged Qin bamboo slip imitation image generation method of claim 1 and a corresponding intact training image.
[0073] S203: Constructing a damaged Qin bamboo slip character image dataset. The damaged Qin bamboo slip character image dataset is combined with the damaged Qin bamboo slip character imitation image dataset to obtain the damaged Qin bamboo slip character image dataset, wherein the damaged Qin bamboo slip character image dataset includes a plurality of damaged Qin bamboo slip character image data, each damaged Qin bamboo slip character image data including a damaged Qin bamboo slip character training image and a corresponding intact training image, and the damaged Qin bamboo slip character training images include real damaged Qin bamboo slip character training images and imitation damaged Qin bamboo slip character training images.
[0074] The purpose of implementing steps S201 to S203 is to aggregate the image data of damaged Qin bamboo slips and mix the real image data of damaged Qin bamboo slips and the imitated image data of damaged Qin bamboo slips into the damaged Qin bamboo slips image data set to facilitate subsequent model training.
[0075] S204: Initialize a generative adversarial inpainting network. The generative adversarial inpainting network includes a inpainting generator and a inpainting discriminator. The inpainting generator is used to output a predicted image of intact Qin bamboo slips based on an input image of damaged Qin bamboo slips. The inpainting discriminator is used to evaluate the authenticity of the intact Qin bamboo slips image input to the inpainting discriminator.
[0076] S205: Training the Generative Adversarial Repair Network. The Generative Adversarial Repair Network is iteratively trained using the training data from the damaged Qin bamboo slips text dataset until training is complete. The basic form of the Generative Adversarial Repair Network remains consistent with the prior art. Under training data from the damaged Qin bamboo slips text dataset, the repair generator and repair discriminator achieve parameter adjustment of the Generative Adversarial Repair Network through mutual adversarial means.
[0077] Specifically, the damaged Qin bamboo slips text image input to the repair generator during the training process is the damaged Qin bamboo slips text training image, and the intact Qin bamboo slips text image input to the repair discriminator during the training process is the predicted intact Qin bamboo slips text image or the intact training image.
[0078] The repair process includes: S206: inputting an unknown damaged Qin bamboo slip text into the generative adversarial repair network, and the predicted intact Qin bamboo slip text image generated by the generator of the generative adversarial repair network is the repaired Qin bamboo slip text image after the unknown damaged Qin bamboo slip text is repaired.
[0079] Example 3: The embodiment of the present invention discloses a restoration generator. The restoration generator of the embodiment of the present invention adopts an encoder-decoder structure, and realizes mapping from low-dimensional features to high-resolution images by integrating residual connections and multi-scale linear attention mechanisms.
[0080] Figure 4 Schematic diagram of the repair generator structure of an embodiment of the present invention.
[0081] Specifically, assuming the input image of damaged Qin bamboo slips is The resolution is 64×64;
[0082] In the encoding stage, the input image first passes through two residual blocks ResBlock1 (the first residual block) and ResBlock2 (the second residual block) and the bottleneck convolution layer conv3 to gradually extract the multi-scale features of the image. Among them, ResBlock1 and ResBlock2 implement spatial downsampling through convolution operations with a stride of 2, gradually reducing the input image resolution from 64×64 to 16×16, while the number of channels is expanded from 1 to 128.
[0083] The encoder design follows the feature pyramid principle of convolutional neural networks, and achieves balanced extraction of global and local features of the image by gradually reducing the spatial resolution and increasing the channel dimension.
[0084] In the decoding stage, the restoration generator gradually restores the spatial resolution of the image through multiple cascaded first upsampling modules Upsample4, second upsampling layer modules Upsample5 and third upsampling modules Upsample6.
[0085] In particular, the repair generator introduces a multi-scale linear attention module (MSLA) in the first two layers of the decoder to enhance the model's contextual perception of the incomplete area, thereby generating reasonable completion content more accurately. That is, the multi-scale linear attention module is introduced in the first upsampling module and the second upsampling module.
[0086] MSLA generates Q (Query), K (Key) and V (Value) by introducing multi-scale information after spatial downsampling, thereby reducing computational complexity and enhancing feature expression capabilities through residual connections.
[0087] Specifically, Q, K, and V are all derived from the input data itself and are vectors generated based on the input features. Q is the query vector, representing the location of the current pixel or local feature. Each Q represents the contextual semantic requirements of a location. K is the key vector, a semantic representation of all pixels in the entire image. It is used to calculate similarity with Q to determine whether a location contributes to Q. V is the value vector, representing the actual content information carried by each location. The attention mechanism assigns weights to V based on the similarity between Q and K, and then performs a weighted sum.
[0088] Finally, the ordinary convolution layer conv7 converts the image into a single-channel grayscale image and outputs the final repaired image.
[0089] Specifically, for the repair generator, except for the output layer whose activation function is Tanh, the activation functions of other layers are ReLU.
[0090] It should be noted that, depending on the degree of resolution restoration of the upsampling layer, a bilinear interpolation layer is generally required between the ordinary convolution layer and the third upsampling module to adjust the resolution of the data, so that the input data and output data of the repair generator have consistent resolution.
[0091] Assume that the input to the generator is a damaged Qin bamboo slip text image, which is a single-channel binary image with a spatial size of 64×64, recorded as ;
[0092] Specifically, the repair generator's specific processing method for input data includes:
[0093] The first layer is a residual block structure, using a 5×5 convolution kernel, a step size of 2, and 32 output channels, reducing the image size to half of the original size;
[0094] The second layer is a residual block structure, using a 3×3 convolution kernel, a stride of 2, and 64 output channels, and the image size is halved again;
[0095] The third layer is the bottleneck convolution layer, which uses a 3×3 convolution kernel, has 128 output channels, and the spatial size remains unchanged;
[0096] The fourth layer consists of an upsampling module and a multi-scale linear attention module (MSLA). Specifically, a bilinear interpolation layer is used to upsample the image to twice its original size. This is followed by a spectral normalization convolution layer with a 3×3 kernel and 64 output channels. Finally, a multi-scale linear attention layer (MSLA) is introduced, which uses dynamic weight allocation to restore high-frequency information.
[0097] The fifth layer structure is the same as the fourth layer structure, consisting of an upsampling module and a multi-scale linear attention module MSLA. The image size is doubled again, and the number of output channels is 32.
[0098] The sixth layer is an upsampling module structure, the image size is doubled again, and the number of output channels is 16;
[0099] A bilinear interpolation layer is used at the connection between the sixth and seventh layers to adjust the image resolution (reduce it by half);
[0100] The seventh layer is a normal convolution layer, which reduces the number of channels to 1 and generates the final single-channel repaired image. .
[0101] Figure 5 2 is a schematic structural diagram of a residual block ResBlock according to an embodiment of the present invention.
[0102] Specifically, the main path of the residual block contains two convolutional layers, which use spectral normalization to stabilize training and batch normalization to accelerate convergence and are activated by ReLU. Before output, the residual block continues to fuse the output data of the main path with the features of the input data.
[0103] Figure 6 2 is a schematic structural diagram of an upsampling module Upsample according to an embodiment of the present invention.
[0104] Specifically, the upsampling module first performs two-fold upsampling through bilinear interpolation, then adjusts the number of channels through 3×3 spectral normalization convolution, and finally cooperates with batch normalization and ReLU to activate and output the data.
[0105] Figure 7 Schematic diagram of the structure of the multi-scale linear attention module MSLA in an embodiment of the present invention.
[0106] Specifically, the multi-scale linear attention module performs average pooling on the input feature map to extract downsampled features, and then projects them in the Q, K, and V directions through convolution respectively, keeping Q at its original scale (projecting after directly performing feature mapping on the original data), using downsampled features for K and V (projecting after average pooling the original data through feature mapping), and performing linear attention calculation (Linear Attention) in a multi-head manner. After normalizing K through softmax, the product of Q and K is calculated and applied to V to capture cross-scale contextual information; finally, the result is restored to the original channel dimension and residually connected with the input to enhance feature expression capabilities.
[0107] Specifically, MSLA is placed after the upsampling stage (the first and second upsampling modules) of the generator. It corresponds to the process of gradually restoring a high-resolution image from a low-resolution image, primarily used to capture long-range dependencies in the input feature map. By modeling the spatial information of the feature map, MSLA ensures that the features of each pixel not only rely on local neighborhood information but also capture global contextual information, thereby improving the model's expressiveness.
[0108] Specifically, for the MSLA module, assuming the input feature map is , , where B represents the batch size, C represents the number of channels, and H×W represents the image spatial resolution.
[0109] Enter Projection is Q, K, V required for multi-head attention:
[0110]
[0111]
[0112]
[0113] Among them, h represents the number of attention heads (head), , is the number of channels per head, is the compression ratio, Indicates that the average pooling operation extracts the multi-scale features after downsampling. In this embodiment, the pooling kernel size is set to 2×2 and stride=2, so the output resolution is .
[0114] In the attention mechanism, are query weight, key weight, and value weight respectively. In the multi-scale linear attention structure of this embodiment, It is a linear projection matrix that maps the input features to the Q, K, V space. These projection matrices are learnable parameters during training and determine the key capabilities of information extraction and reconstruction in the attention mechanism.
[0115] The traditional attention computation complexity is ,in , the calculation formula is as follows:
[0116]
[0117] The computational complexity of this model is The model and It is through the input The downsampled features obtained after spatial pooling are then linearly transformed, so the computational complexity is greatly reduced.
[0118] First of all, Perform softmax normalization in the channel dimension and spatial dimension:
[0119]
[0120] Next, calculate the attention output:
[0121]
[0122] The softmax in this model is a direct Normalization is performed to avoid large-scale N×N similarity calculations, thereby reducing computational complexity. At the same time, downsampling is used to achieve multi-scale modeling, taking into account both global receptive field and computational efficiency.
[0123] Specifically, the MSLA processing adopted by the first upsampling module mainly processes low-level features and enhances the recovery ability of coarse structures.
[0124] Specifically, the MSLA adopted by the second upsampling module mainly processes fine features to help restore details and accurate stroke information.
[0125] Example 4:
[0126] Figure 8 Schematic diagram of the repair discriminator structure of an embodiment of the present invention.
[0127] Specifically, the restoration discriminator of the embodiment of the present invention is a typical DCGAN discriminator, which includes multiple convolutional layers and activation functions.
[0128] Specifically, the predicted intact Qin bamboo slips text images generated by the repair generator will be (or a complete training image ) is input to the restoration discriminator, which gradually extracts features from the image through four convolutional layers (downsample layers) and reduces the spatial resolution of the image. Finally, the convolutional layer outputs a single value to evaluate the authenticity of the image. Finally, the Sigmoid activation function is used to map the output to the range of [0, 1] to determine the authenticity probability of the image.
[0129] Figure 9 2 is a schematic diagram of the structure of the downsample layer according to an embodiment of the present invention.
[0130] 4×4 convolution kernels with a stride of 2 are used for progressive downsampling, halving the resolution at each layer and increasing the number of channels from 64 to 512. Each layer is activated with a LeakyReLU after the convolution kernel processing. The fifth layer uses a 3×3 convolution to compress the feature map into a 1×1 output, and a sigmoid activation to generate the discriminant probability. Spectral normalization is used for all convolutional layers to stabilize training.
[0131] Throughout the training process, the inpainting generator and the inpainting discriminator are alternately optimized, and the model parameters are refined through training. The training process can be summarized as an adversarial game between the inpainting generator and the inpainting discriminator. The inpainting discriminator strives to distinguish between real and generated images, while the inpainting generator strives to generate realistic images, thereby "deceiving" the discriminator.
[0132] The loss of the inpainting discriminator consists of two parts: the binary classification loss of the intact training images And the binary classification loss of predicting intact Qin bamboo slip text images .
[0133]
[0134]
[0135] in, and Both are binary cross entropy losses. ,
[0136] The total loss of repair discrimination is:
[0137]
[0138] The goal of the restoration generator is to minimize the judgment error of the restoration discriminator on the generated image, making the generated image more like the real image.
[0139] The loss function of the repair generator includes the reconstruction loss , against loss and perceptual loss .
[0140]
[0141] Reconstruction losses Generate an image With the target image The mean square error loss between .
[0142]
[0143] is the generation adversarial loss, which the generator hopes to maximize , making its output close to 1.
[0144]
[0145] is the definition of perceptual loss, where , , The feature maps are The number of channels, height and width of the layer. If only the features of a single layer are selected for comparison, the perceptual loss can be simplified to:
[0146]
[0147] The total loss of the generator is:
[0148]
[0149] in, and Represents the weight factor, set , .
[0150] It should be noted that in the entire conditional repair generation network, the basic component structures such as convolutional layers belong to the existing technology. The basic component structures used in different modules are the same, and the actual function is determined by the setting position of the basic component in the model.
[0151] Specifically, in each training step, the repair discriminator accepts predicted intact Qin bamboo slips text images and intact training images for training. The repair discriminator is optimized by maximizing its prediction probability for intact training images and minimizing its prediction probability for predicted intact Qin bamboo slips text images. The repair generator is optimized according to its loss. The repair generator is trained by minimizing the difference between its predicted intact Qin bamboo slips text images and intact training images and trying its best to make the repair discriminator believe that the predicted intact Qin bamboo slips text images it outputs are intact training images.
[0152] Specifically, after each iteration, the weights of the repair generator and the repair discriminator are updated through the back-propagation algorithm, and the optimizer is Adam.
[0153] In practical applications, the damaged Qin bamboo slips text image data in the damaged Qin bamboo slips text image data set can be divided into data for training and data for testing; after the training of the repair generator and the repair discriminator is completed, the data for testing can be used to test the trained repair generator, and the PSNR and SSIM can be calculated according to the output results to evaluate the repair performance of the repair generator of the embodiment of the present invention.
[0154] Embodiment 5:
[0155] Figure 10 This is a schematic diagram comparing the physical image of Qin bamboo slips characters and the binary image of Qin bamboo slips characters according to an embodiment of the present invention.
[0156] Specifically, multiple images of Qin bamboo slips characters are extracted from data that are not involved in training and testing. In order to meet the model data form requirements of the embodiment of the present invention (the input requirement is a binary image), it can be processed as follows: the grayscale background of the Qin bamboo slips characters image is removed and converted into a binary image of Qin bamboo slips characters. Specifically, the background removal operation first converts the input image into a grayscale image, reads the grayscale image, converts it into a single-channel image, and sets the threshold through the binarization operation. (Can be set to 128 according to actual needs), pixel value is greater than or equal to Set the pixel value to 255, which is less than The pixel value is set to 0 to achieve the purpose of background elimination.
[0157] Refer to the attached figure Figure 10, after binarization, the corresponding Qin bamboo slip text entity image forms a corresponding binary image of the Qin bamboo slip text. It should be noted that for the sake of clarity of illustration, Figure 10 The sampled Qin bamboo slip text entity image is a Qin bamboo slip text entity image with a relatively complete physical structure.
[0158] Figure 11 This is a schematic diagram of the repair example of the repair generator in the embodiment of the present invention. Among them, the first column is the unknown damaged Qin bamboo slip text, the second column is the repaired Qin bamboo slip text image, and the third column is the manually repaired Qin bamboo slip text image (after binarization). Refer to the attached drawing Figure 11 It is shown that for different configurations of the character "ke" and different damage situations, the repaired Qin bamboo slip text image generated by the real-time repair generator of the present invention has a high similarity with the manually repaired Qin bamboo slip text image, and can well replace manual labor for automated Qin bamboo slip text repair work.
[0159] In summary, the present invention provides a method for generating a forged image of damaged Qin bamboo slip text and an image repair method. The method for generating a forged image of damaged Qin bamboo slip text uses all types of image data in the same period as the reference object for extracting the damage mask, extracts the damage mask through a generative adversarial damage network, and then processes the intact Qin bamboo slip text image with the extracted damage mask to obtain a highly realistic damaged Qin bamboo slip text image. This method provides a method that can effectively expand the number of damaged ancient character samples and has good practicability in machine learning applications in related fields.
[0160] The above has introduced in detail a method for generating a forged image of damaged Qin bamboo slip text and an image repair method provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for generating an imitation image of damaged Qin bamboo slips, characterized in that: include Constructing a damaged image dataset, the damaged image dataset including a plurality of damaged training data, each damaged training data including a damaged training image and a corresponding intact training image, the intact training images including intact training images of Qin bamboo slips characters and intact training images of non-Qin bamboo slips characters, and the damaged training images including damaged training images of Qin bamboo slips characters and damaged training images of non-Qin bamboo slips characters; Initializing a generative adversarial damage network, which includes a mask generator, a damage generator, and a damage discriminator. The mask generator is configured to output a predicted damage mask based on an input intact image. The damage generator is configured to process the corresponding intact image according to the predicted damage mask to obtain a predicted damaged image. The damage discriminator is configured to evaluate the authenticity of the input damaged image according to a preset damage evaluation index. Training a generative adversarial network, iteratively training the generative adversarial network using the damaged training data in the damaged image dataset until the training is completed; Constructing a mask dataset, sequentially importing intact training images from the damaged image dataset into a trained generative adversarial network, causing a mask generator in the generative adversarial network to sequentially generate corresponding predicted damage masks, and storing all predicted damage masks generated by the mask generator in the mask dataset; The imitation image of damaged Qin bamboo slips is generated by randomly extracting a predicted damage mask from the mask data set to process any intact training image of Qin bamboo slips in the damaged image data set to obtain the corresponding imitation damaged training image, which is the required imitation image of damaged Qin bamboo slips.
2. The method for generating a replica image of damaged Qin bamboo slips as claimed in claim 1, characterized in that: The damaged training image has a preset fixed resolution.
3. The method for generating a replica image of damaged Qin bamboo slips as claimed in claim 1, characterized in that: The damaged training image and the intact training image are both binarized images.
4. A method for repairing damaged Qin bamboo slips, including a training process and a repair process, characterized in that: The training process includes: Constructing a damaged Qin bamboo slips character real image dataset, wherein the damaged Qin bamboo slips character real image dataset includes a plurality of damaged Qin bamboo slips character real image data, each of which includes a damaged Qin bamboo slips character real image training image and a corresponding intact training image; Constructing a damaged Qin bamboo slips imitation image dataset, wherein the damaged Qin bamboo slips imitation image dataset includes a plurality of damaged Qin bamboo slips imitation image data, each of which includes a damaged Qin bamboo slips imitation image generated by the damaged Qin bamboo slips imitation image generation method according to any one of claims 1 to 3 and a corresponding intact training image; Constructing a damaged Qin bamboo slip character image dataset, merging the damaged Qin bamboo slip character real image dataset and the damaged Qin bamboo slip character imitation image dataset to obtain the damaged Qin bamboo slip character image dataset, wherein the damaged Qin bamboo slip character image dataset includes a plurality of damaged Qin bamboo slip character image data, each damaged Qin bamboo slip character image data includes a damaged Qin bamboo slip character training image and a corresponding intact training image, and the types of the damaged Qin bamboo slip character training images include the damaged Qin bamboo slip character real training images and the damaged Qin bamboo slip character imitation training images; Initializing a generative adversarial restoration network, which includes a restoration generator and a restoration discriminator. The restoration generator is used to output a predicted intact Qin bamboo slip text image based on an input damaged Qin bamboo slip text image, and the restoration discriminator is used to evaluate the authenticity of the intact Qin bamboo slip text image input to the restoration discriminator. Train a generative adversarial repair network, and iteratively train the generative adversarial repair network using the training data in the damaged Qin bamboo slips text dataset until the training is completed.
5. The method for repairing damaged Qin bamboo slips as claimed in claim 4, characterized in that: The repair process includes: An unknown damaged Qin bamboo slip text is input into the generative adversarial repair network, and the predicted intact Qin bamboo slip text image generated by the generator of the generative adversarial repair network is the repaired Qin bamboo slip text image after the unknown damaged Qin bamboo slip text has been repaired.
6. The method for repairing damaged Qin bamboo slips as claimed in claim 4, characterized in that: The restoration generator includes a first residual block, a second residual block, a bottleneck convolution layer, a first upsampling module, a second upsampling module, a third upsampling module, a bilinear interpolation layer and a normal convolution layer, which are linked in sequence.
7. The method for repairing damaged Qin bamboo slips as claimed in claim 6, characterized in that: The first upsampling module and the second upsampling module are respectively provided with a multi-scale linear attention module.
8. The method for repairing damaged Qin bamboo slips as claimed in claim 6 or 7, characterized in that: The loss function of the inpainting generator includes reconstruction loss, adversarial loss and perceptual loss.
9. The method for repairing damaged Qin bamboo slips as claimed in claim 4, characterized in that: The restoration discriminator is a DCGAN discriminator.
10. The method for repairing damaged Qin bamboo slips as claimed in claim 7, characterized in that: The loss function of the restoration discriminator includes the binary classification loss of the predicted intact Qin bamboo slip text image and the binary classification loss of the intact training image.
Citation Information
Patent Citations
Hyperspectral ancient painting detection and recognition method based on deep learning
CN111291675A
Construction method and device of classification model for evaluating implant stability based on CBCT image data
CN115512167A
Image restoration method
CN116051407A
Qinxi character restoration method based on conditional generative adversarial network
CN116681604A
Urban space distortion image restoration evaluation method
CN117197730A