A damaged qin dynasty bamboo slip character copy image generation method and image restoration method
By generating a damage mask using image data from the same period through a generative adversarial damage network, the problem of large differences between the simulated ancient text images and the real images in the existing technology is solved. This achieves highly realistic generation and restoration of damaged Qin bamboo slip text images and improves the restoration effect of machine learning.
Patent Information
- Application Number
- CN202510761491.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-06-09
AI Technical Summary
Existing technologies using machine learning to restore ancient text images often result in poor restoration outcomes due to significant differences between the simulated training samples and the actual damaged ancient text images.
By constructing a generative adversarial network for damage analysis, and using all types of image data from the same period as reference objects for extracting damage masks, highly realistic images of damaged Qin bamboo slips are generated, thus expanding the number of damaged ancient character samples.
It improves the simulation and effectiveness of image restoration of damaged ancient characters and enhances the application effect of machine learning models in ancient character restoration.
Smart Images

Figure CN120672888B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of ancient script restoration, specifically to a method for generating images of damaged Qin Dynasty bamboo slips and a method for restoring images of damaged Qin Dynasty bamboo slips. Background Technology
[0002] With the development of science and technology, many methods for image restoration of ancient characters using machine learning techniques such as neural networks and generative adversarial networks have emerged. Image restoration of ancient characters using machine learning methods requires a certain amount of training data. In order to obtain a large amount of training data, training samples can be generated through imitation during the data acquisition process.
[0003] The publicly available document CN116681604A discloses methods for generating images of damaged ancient characters, including image masking, image rotation, image cropping, image scaling, salt-and-pepper noise, and Gaussian noise. In practical applications, these methods are essentially formulaic processing techniques. The damaged ancient character images generated using these methods differ significantly from genuine damaged ancient character images. If such simulated training samples are used for machine learning, the resulting model will not be able to effectively repair genuine damaged ancient character images.
[0004] Therefore, a more reasonable method for replicating damaged ancient scripts is needed to meet the requirements of machine learning in related fields. Summary of the Invention
[0005] This invention provides a method for generating and restoring images of damaged Qin bamboo slip characters. This method utilizes all types of image data from the same period as reference objects for extracting the damage mask. It extracts the damage mask by generating an adversarial damage network, and then uses the extracted damage mask to process intact Qin bamboo slip character images, thereby obtaining highly realistic images of damaged Qin bamboo slip characters. This method provides an effective way to expand the number of damaged ancient character samples and has good practicality in machine learning applications in related fields.
[0006] Accordingly, the present invention provides a method for generating images of damaged Qin bamboo slips, including...
[0007] A damaged image dataset is constructed, which includes several damaged training data. Each damaged training data includes a damaged training image and a corresponding intact training image. The intact training images include intact training images of Qin bamboo slips and intact training images of non-Qin bamboo slips. The damaged training images include damaged training images of Qin bamboo slips and damaged training images of non-Qin bamboo slips.
[0008] An initialization of a generative adversarial network is performed. The generative adversarial network includes a mask generator, a damage generator, and a damage discriminator. The mask generator is used to output a predicted damage mask based on the input intact image. The damage generator is used to process the corresponding intact image based on the predicted damage mask to obtain a predicted damaged image. The damage discriminator is used to evaluate the authenticity of the input damaged image based on a preset damage evaluation index.
[0009] The generative adversarial network is trained by iteratively training the network using damaged training data from the damaged image dataset until training is complete.
[0010] Construct a mask dataset, and sequentially import the intact training images from the damaged image dataset into the trained generative adversarial destruction network. The mask generator in the generative adversarial destruction network sequentially generates corresponding predicted destruction masks, and all predicted destruction masks generated by the mask generator are stored in the mask dataset.
[0011] The process involves generating a simulated image of Qin bamboo slip characters by randomly extracting a predicted damage mask from the mask dataset and processing any intact Qin bamboo slip character training image from the damaged image dataset to obtain a corresponding simulated damaged training image. This simulated damaged training image is the desired simulated image of Qin bamboo slip characters.
[0012] In an optional implementation, the damaged training image has a preset fixed resolution.
[0013] In an optional implementation, both the damaged training image and the intact training image are binarized images.
[0014] Accordingly, the present invention provides a method for restoring damaged Qin Dynasty bamboo slip text images, including a training process and a restoration process, wherein the training process includes:
[0015] A dataset of real images of damaged Qin bamboo slips is constructed. The dataset of real images of damaged Qin bamboo slips includes several real images of damaged Qin bamboo slips. Each real image of damaged Qin bamboo slips includes a training real image of the damaged Qin bamboo slips and a corresponding intact training image.
[0016] A dataset of images of Qin bamboo slips with inscriptions is constructed. The dataset includes several sets of images of Qin bamboo slips with inscriptions. Each set of images includes an image of Qin bamboo slips with inscriptions generated by the image generation method and a corresponding intact training image.
[0017] A dataset of images of damaged Qin bamboo slips is constructed by merging the dataset of real images of damaged Qin bamboo slips with the dataset of simulated images of damaged Qin bamboo slips. The dataset of images of damaged Qin bamboo slips includes several sets of images of damaged Qin bamboo slips. Each set of images includes a training image of a damaged Qin bamboo slip and a corresponding intact training image. The types of training images of damaged Qin bamboo slips include real training images of damaged Qin bamboo slips and simulated training images of damaged Qin bamboo slips.
[0018] An initialization generative adversarial repair network is performed, which includes a repair generator and a repair discriminator. The repair generator is used to output a predicted image of intact Qin bamboo slips based on the input image of damaged Qin bamboo slips. The repair discriminator is used to evaluate the authenticity of the image of intact Qin bamboo slips input to the repair discriminator.
[0019] The generative adversarial repair network is trained by iteratively training the network using training data from the damaged Qin bamboo slips text dataset until training is complete.
[0020] In an optional implementation, the repair process includes:
[0021] An unknown damaged Qin bamboo slip character is input into the generative adversarial repair network. The generator of the generative adversarial repair network generates a predicted intact Qin bamboo slip character image, which is the repaired Qin bamboo slip character image after the unknown damaged Qin bamboo slip character has been repaired.
[0022] In an optional implementation, the repair generator includes a first residual block, a second residual block, a bottleneck convolutional layer, a first upsampling module, a second upsampling module, a third upsampling module, a bilinear interpolation layer, and a normal convolutional layer, which are linked in sequence.
[0023] In an optional implementation, the first upsampling module and the second upsampling module are each respectively provided with a multi-scale linear attention module.
[0024] In an optional implementation, the loss function of the repair generator includes reconstruction loss, adversarial loss, and perceptual loss.
[0025] In an optional implementation, the repair discriminator is a DCGAN discriminator.
[0026] In an optional implementation, the loss function of the repair discriminator includes the binary classification loss of the predicted intact Qin bamboo slip text image and the binary classification loss of the intact training image.
[0027] This invention provides a method for generating and restoring images of damaged Qin bamboo slip characters. The method utilizes all types of image data from the same period as reference objects for extracting the damage mask. It extracts the damage mask through an adversarial damage network, and then uses the extracted damage mask to process intact Qin bamboo slip character images, thereby obtaining highly realistic images of damaged Qin bamboo slip characters. This method provides an effective way to expand the number of damaged ancient character samples and has good practicality in machine learning applications in related fields. Attached Figure Description
[0028] Figure 1 This is a flowchart of the method for generating images of damaged Qin Dynasty bamboo slips according to Embodiment 1 of the present invention.
[0029] Figure 2 This is an example diagram of a damaged training data set from Embodiment 1 of the present invention.
[0030] Figure 3 This is a flowchart of the method for restoring damaged Qin Dynasty bamboo slip text images according to an embodiment of the present invention.
[0031] Figure 4 This is a schematic diagram of the repair generator structure according to Embodiment 3 of the present invention.
[0032] Figure 5 This is a schematic diagram of the structure of the residual block ResBlock in Embodiment 3 of the present invention.
[0033] Figure 6 This is a schematic diagram of the Upsample module in Embodiment 3 of the present invention.
[0034] Figure 7 This is a schematic diagram of the structure of the multi-scale linear attention module (MSLA) in Embodiment 3 of the present invention.
[0035] Figure 8 This is a schematic diagram of the repair discriminator structure in Embodiment 4 of the present invention.
[0036] Figure 9 This is a schematic diagram of the downsample layer structure in Embodiment 4 of the present invention.
[0037] Figure 10 This is a schematic diagram comparing the physical image of Qin bamboo slip characters and the binarized image of Qin bamboo slip characters in Embodiment 5 of the present invention.
[0038] Figure 11 This is a schematic diagram of a repair generator repair example according to Embodiment 5 of the present invention. Detailed Implementation
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0040] Embodiment 1
[0041] Figure 1 It is a flowchart of a method for generating a forged image of damaged Qin bamboo slips characters according to an embodiment of the present invention.
[0042] The embodiments of the present invention disclose a method for generating a forged image of damaged Qin bamboo slips characters, including
[0043] S101: Construct a damaged image data set;
[0044] The damaged image data set includes a number of damaged training data. Each piece of damaged training data includes a damaged training image and a corresponding intact training image. The damaged training image and the intact training image have the same resolution. The damaged training image includes a damaged training image of Qin bamboo slips characters and a damaged training image of non-Qin bamboo slips characters. The intact training image includes an intact training image of Qin bamboo slips characters and an intact training image of non-Qin bamboo slips characters;
[0045] Figure 2 It is an example diagram of a piece of damaged training data according to an embodiment of the present invention. Specifically, the meaning of this piece of damaged training data is the character 'ke'. The intact training image of Qin bamboo slips characters is a clear binary image of the character 'ke', and the damaged training image is a binary image of the character 'ke' with defects.
[0046] It should be noted that in this step, the damaged training image includes a damaged training image of Qin bamboo slips characters and a damaged training image of non-Qin bamboo slips characters, and the intact training image includes an intact training image of Qin bamboo slips characters and an intact training image of non-Qin bamboo slips characters. That is, although the embodiments of the present invention need to generate a forged image of damaged Qin bamboo slips characters, when collecting damaged training data, the images of other works in the same period can be collected in the same data form. The internal logic of this implementation means is that the image of damaged Qin bamboo slips characters is essentially an image, and the damaged traces are substantially the same as those of non-Qin bamboo slips characters. The purpose of the embodiments of the present invention is to extract the damaged mask, and the damaged mask reflects the damaged characteristics rather than the characteristics of the characters. Therefore, in the collection stage of training data, the damaged training image includes a damaged training image of Qin bamboo slips characters and a damaged training image of non-Qin bamboo slips characters, and the intact training image includes an intact training image of Qin bamboo slips characters and an intact training image of non-Qin bamboo slips characters.
[0047] S102: Initialize and generate an adversarial network;
[0048] The generative adversarial damage network includes a mask generator, a damage generator, and a damage discriminator.
[0049] The mask generator is used to output a predicted damaged mask based on the input intact image. Specifically, assume the input intact training image is I. good After processing by the mask generator, a predicted damage mask M is output. It should be noted that since the data format used in this embodiment is a binary image, which does not contain blurred content, the final damage mask is actually an occlusion mask. This determines the definition formula of the damage generator. If the data format is a grayscale image or a color image, a blurred mask is more likely. When using a blurred mask, the damage generator formula needs to include a parameter regarding the mask's effectiveness.
[0050] The damage generator is used to process the intact training image according to the predicted damage mask to obtain a predicted damage image. Specifically, the predicted damage image... Where the symbol ⊙ represents the Hadamard product, Interpreted as random masking noise. Indicates to masking The mask is applied directly to the intact image to obtain the first intermediate image after masking; then it is overlaid on the first intermediate image. This is based on noise. Random masking is performed in areas where the mask is not applied. It can be understood as a compensation parameter that can be adjusted during training, so that the model has a certain degree of flexibility during training, in order to avoid the situation where the input result and the output result are completely equal;
[0051] The damage discriminator is used to evaluate the authenticity of the input predicted damage image or the damage training image according to a preset damage evaluation index.
[0052] Specifically, the damage generator can adopt the U-Net structure, and the damage discriminator can adopt the PatchGAN structure.
[0053] Specifically, the damaged generator adopts a U-Net structure, which includes a damaged encoder and a damaged decoder. The damaged encoder primarily executes a downsampling path, consisting of multiple convolutional layers that progressively extract high-level semantic features and reduce the spatial resolution of the input intact image. The damaged decoder primarily executes an upsampling path, consisting of multiple transposed convolutional layers (used for upsampling) that progressively restore the spatial resolution of the image and generate the required mask. Specifically, the number of convolutional layers used for downsampling in the damaged encoder and the number of transposed convolutional layers used for upsampling in the damaged decoder are strictly corresponding, and a skip link needs to be set between the damaged encoder and the damaged decoder. Specifically, the feature map generated by a certain convolutional layer used for upsampling in the encoder is skipped to the corresponding layer (with the same resolution) of the device convolutional layer in the decoder. This achieves the following: passing low-level detail information (details lost by the encoder during compression are directly passed to the decoder through skip links); mitigating gradient vanishing (providing a shortcut for gradients to flow directly from the deep layers of the decoder to the shallow layers of the encoder); and fusing multi-scale information (merging features of different scales to ensure the generated image has global structural consistency).
[0054] Specifically, the damage generator is divided into layers according to the number of downsampling operations, generally requiring 4 to 8 layers. Since the intact image in this embodiment has low complexity and low resolution (generally 128x128), the number of layers is generally set to 4 to 5. In each layer, a regular convolution with a 3x3 kernel or a 4x4 kernel and a stride of 1 is set for feature extraction, and a LeakyReLU or ReLU is set for feature activation. Then, downsampling is performed through pooling (parameters set to 2x2 pooling, stride of 2) or stride convolution (parameters set to 4x4 convolution, stride of 2). After the downsampling layers of the last layer are completed, 2 to 3 bottleneck convolutional layers are generally set for the most abstract feature extraction.
[0055] Specifically, the damaged decoder is divided into layers according to the number of upsampling operations, with the number of layers being the same as that of the damaged generator. In each layer, a transposed convolution (with parameters set to 4x4 kernels and a stride of 2) is generally used for upsampling. Specifically, the skip link point between the damaged encoder and the damaged decoder can be set after the transposed convolution of any layer (the feature map resolution of the corresponding link of the damaged encoder must be consistent with the resolution of the output of the transposed convolution). Generally, the skip link is set after the transposed convolution of the last layer. Specifically, the fusion method at the skip link is a splicing method (addition method can be used). Before the mask output, after ordinary convolution processing and activation function activation, the required mask is generated by normalization function (generally Tanh). The purpose of normalization is that the mask itself has normalized features (black and white features), and normalization can well coordinate the format of the output data.
[0056] Specifically, the damage discriminator adopts the PatchGAN structure. PatchGAN is a discriminator architecture specifically designed for image generation tasks (especially image-to-image translation). Its core idea is not to directly judge the authenticity of the entire image, but to independently judge the local regions (patch, receptive field of view) of the image, and finally output a statistical graph or statistical matrix about the "local authenticity" of each receptive field of view. Each element in the statistical graph or statistical matrix represents the probability that the corresponding local region (patch) in the input image is a real image.
[0057] Specifically, the receptive field of the damage discriminator (i.e., the size of the input image region "seen" by each output neuron) can be 1x1, 16x16, 30x30, etc. The output of the damage discriminator is a statistical matrix, but during training, the average value of all output units is usually taken as the final discrimination loss of the entire image.
[0058] Specifically, the input to the damage discriminator is a stitched image of the predicted damage image and the damage training image. The damage discriminator will calculate the similarity (true probability) of each receptive field based on the receptive field and output a statistical matrix. The average value of the statistical matrix is used as the damage evaluation index. The loss function can be a superposition of adversarial loss (cGAN loss) and L1 regularization loss.
[0059] S103: Training to generate adversarial compromise networks;
[0060] The generative adversarial network is iteratively trained using the damaged training data in the damaged image dataset until training is complete.
[0061] Specifically, the generative adversarial network of this invention is not fundamentally different in form from the generative adversarial network of the prior art. The innovation of the generative adversarial network of this invention lies mainly in the structural design of the generative adversarial network (the addition of a mask generator) and the selection of damage training data (the damage training images include Qin bamboo slip text damage training images and non-Qin bamboo slip text damage training images). Therefore, the structure of the damage generator and the damage discriminator can be designed based on the prior art, and their training process can also be carried out according to the prior art. This invention does not impose any additional limitations.
[0062] S104: Construct the mask dataset;
[0063] The intact training images of the damaged training data in the damaged image dataset are sequentially imported into the trained generative adversarial network. The mask generator in the generative adversarial network sequentially generates the corresponding predicted damage masks, and the predicted damage masks are stored in a mask dataset.
[0064] After the generative adversarial network is trained, when a complete training image is input into the generative adversarial network, the mask generator will export the corresponding mask; traversing the input complete training images, the mask generator will export the corresponding number of masks; the masks exported by the mask generator are stored in a mask dataset.
[0065] S105: Generation of a replica image of the damaged Qin bamboo slips. A predicted damage mask is randomly extracted from the mask dataset and processed on any intact training image of the Qin bamboo slips in the damaged image dataset to obtain a corresponding replica damaged training image. The replica damaged training image is the required replica image of the damaged Qin bamboo slips.
[0066] The masks in the mask dataset represent the current state of damage to physical objects from the same period. Therefore, this state of damage is not actually corresponding; that is, although the mask is extracted from a corresponding intact training image, any physical object from the same period may have the corresponding state of damage. Therefore, in this step, by using different masks to traverse different intact training images of Qin bamboo slips, a large number of imitation images of damaged Qin bamboo slips can be generated. The imitation images of damaged Qin bamboo slips generated in this way have high simulation accuracy, and the model trained on these imitation images can better fit the actual situation in applications.
[0067] Example 2
[0068] Figure 3 This is a flowchart of the method for restoring damaged Qin Dynasty bamboo slip text images according to an embodiment of the present invention.
[0069] This invention provides a method for restoring damaged Qin Dynasty bamboo slip text images, including a training process and a restoration process.
[0070] The training process includes steps S201 to S205.
[0071] S201: Construct a dataset of real images of damaged Qin bamboo slips. The dataset of real images of damaged Qin bamboo slips includes several sets of real images of damaged Qin bamboo slips. Each set of real images of damaged Qin bamboo slips includes a training real image of the damaged Qin bamboo slip and a corresponding intact training image.
[0072] S202: Construct a dataset of images replicating damaged Qin bamboo slip characters. The dataset includes several sets of images replicating damaged Qin bamboo slip characters, each set including an image replicating damaged Qin bamboo slip characters generated by the aforementioned image generation method and a corresponding intact training image.
[0073] S203: Construct a dataset of images of damaged Qin bamboo slips. The dataset of real images of damaged Qin bamboo slips is merged with the dataset of simulated images of damaged Qin bamboo slips to obtain the dataset of images of damaged Qin bamboo slips. The dataset of images of damaged Qin bamboo slips includes several sets of images of damaged Qin bamboo slips. Each set of images includes one training image of a damaged Qin bamboo slip and a corresponding intact training image. The types of training images of damaged Qin bamboo slips include real training images of damaged Qin bamboo slips and simulated training images of damaged Qin bamboo slips.
[0074] The purpose of steps S201 to S203 is to summarize the image data of the damaged Qin bamboo slips, and to mix the real image data of the damaged Qin bamboo slips and the imitation image data of the damaged Qin bamboo slips into the image dataset of the damaged Qin bamboo slips, so as to facilitate subsequent model training.
[0075] S204: Initialize the generative adversarial repair network. The generative adversarial repair network includes a repair generator and a repair discriminator. The repair generator is used to output a predicted image of intact Qin bamboo slips based on the input image of damaged Qin bamboo slips. The repair discriminator is used to evaluate the authenticity of the image of intact Qin bamboo slips input to the repair discriminator.
[0076] S205: Training the Generative Adversarial Repair Network. The generative adversarial repair network is iteratively trained using training data from the damaged Qin bamboo slips text dataset until training is complete. The basic form of the generative adversarial repair network remains consistent with existing technologies. Under training with the training data from the damaged Qin bamboo slips text dataset, the repair generator and repair discriminator adjust the parameters of the generative adversarial repair network through mutual adversarial interaction.
[0077] Specifically, the damaged Qin bamboo slip text image input to the repair generator is the damaged Qin bamboo slip text training image during the training process, and the intact Qin bamboo slip text image input to the repair discriminator is the predicted intact Qin bamboo slip text image or the intact training image during the training process.
[0078] The repair process includes: S206: Inputting an unknown damaged Qin bamboo slip character into the generative adversarial repair network, wherein the generator of the generative adversarial repair network generates a predicted intact Qin bamboo slip character image, which is the repaired Qin bamboo slip character image after the unknown damaged Qin bamboo slip character has been repaired.
[0079] Example 3: This embodiment of the invention discloses a repair generator. The repair generator of this embodiment adopts an encoder-decoder structure and realizes the mapping from low-dimensional features to high-resolution images by incorporating residual connections and multi-scale linear attention mechanisms.
[0080] Figure 4 This is a schematic diagram of the repair generator structure according to an embodiment of the present invention.
[0081] Specifically, assuming the input is a damaged Qin Dynasty bamboo slip image. The resolution is 64×64;
[0082] During the encoding stage, the input image first passes through two residual blocks, ResBlock1 (first residual block) and ResBlock2 (second residual block), and the bottleneck convolutional layer conv3 to gradually extract multi-scale features of the image. ResBlock1 and ResBlock2 achieve spatial downsampling through convolution operations with a stride of 2, gradually reducing the resolution of the input image from 64×64 to 16×16, while expanding the number of channels from 1 to 128.
[0083] The encoder design follows the feature pyramid principle of convolutional neural networks, achieving a balanced extraction of global and local features of the image by gradually reducing spatial resolution and increasing channel dimensions.
[0084] During the decoding stage, the repair generator gradually restores the spatial resolution of the image through multiple cascaded upsampling modules: Upsample4 (first upsampling layer), Upsample5 (second upsampling layer), and Upsample6 (third upsampling layer).
[0085] In particular, the repair generator introduces a multi-scale linear attention (MSLA) module in the first two layers of the decoder to enhance the model's ability to perceive the context of the missing region, thereby generating reasonable completion content more accurately. That is, the multi-scale linear attention module is introduced in the first upsampling module and the second upsampling module.
[0086] MSLA generates Q (Query), K (Key), and V (Value) by introducing multi-scale information after spatial downsampling, thereby reducing computational complexity and enhancing feature representation capabilities through residual connections.
[0087] Specifically, Q, K, and V all originate from the input data itself, which are vectors generated based on the input features. Q is the query vector, representing the position of the current pixel or local feature; each Q represents the contextual semantic requirement of a position. K is the key vector, the semantic representation of all pixels in the entire image, used to calculate similarity with Q, thereby determining whether a position is helpful to Q. V is the value vector, representing the actual content information carried by each position. The attention mechanism assigns weights to V based on the similarity between Q and K, and then performs a weighted summation.
[0088] Finally, the ordinary convolutional layer conv7 converts the image into a single-channel grayscale image, outputting the final repaired image.
[0089] Specifically, for the repair generator, except for the output layer which uses Tanh activation function, all other layers use ReLU activation function.
[0090] It is important to note that, depending on the resolution restoration level of the upsampling layer, a bilinear interpolation layer is generally required between the ordinary convolutional layer and the third upsampling module to adjust the data resolution, so that the input and output data of the repair generator have consistent resolution.
[0091] Suppose the input to the generator is a damaged Qin Dynasty bamboo slip image, which is a single-channel binarized image with a spatial size of 64×64, denoted as . ;
[0092] Specifically, the repair generator's methods for processing input data include:
[0093] The first layer is a residual block structure, using a 5×5 convolution kernel with a stride of 2 and 32 output channels, reducing the image size to half of the original.
[0094] The second layer is a residual block structure, using a 3×3 convolution kernel with a stride of 2, and the number of output channels is 64. The image size is halved again.
[0095] The third layer is the bottleneck convolutional layer, which uses a 3×3 convolutional kernel and has 128 output channels, while maintaining the same spatial size.
[0096] The fourth layer consists of an upsampling module and a multi-scale linear attention module (MSLA). Specifically, upsampling is performed through a bilinear interpolation layer to double the image size, followed by a spectral normalization convolutional layer with a 3×3 kernel and 64 output channels. Finally, the multi-scale linear attention layer (MSLA) is introduced, which recovers high-frequency information through dynamic weight allocation.
[0097] The fifth layer structure is the same as the fourth layer structure, consisting of an upsampling module and a multi-scale linear attention module (MSLA). The image size is magnified twice again, and the number of output channels is 32.
[0098] The sixth layer is an upsampling module structure, which doubles the image size again and outputs 16 channels.
[0099] At the junction of the sixth and seventh layers, a bilinear interpolation layer is used to adjust the image resolution (reducing it by half).
[0100] The seventh layer is a regular convolutional layer, which reduces the number of channels to 1, generating the final single-channel restored image. .
[0101] Figure 5 This is a schematic diagram of the structure of the residual block ResBlock according to an embodiment of the present invention.
[0102] Specifically, the residual block main path contains two convolutional layers, which are trained using spectral normalization and accelerated by batch normalization, respectively, and activated by ReLU. Before output, the residual block further fuses and links the output data of the main path with the features of the input data.
[0103] Figure 6 This is a schematic diagram of the Upsample module in an embodiment of the present invention.
[0104] Specifically, the upsampling module first performs double upsampling through bilinear interpolation, then adjusts the number of channels through 3×3 spectral normalization convolution, and finally activates and outputs the data in conjunction with batch normalization and ReLU.
[0105] Figure 7 This is a schematic diagram of the structure of the multi-scale linear attention module (MSLA) according to an embodiment of the present invention.
[0106] Specifically, the multi-scale linear attention module performs average pooling on the input feature map to extract downsampled features, and then projects them in the Q, K, and V directions through convolution. Q is kept at its original scale (projection is performed after direct feature mapping of the original data), while K and V use downsampled features (projection is performed after average pooling of the original data through feature mapping). Linear attention is calculated in a multi-head manner. After normalizing K with softmax, the product of Q and K is calculated and applied to V to capture cross-scale contextual information. Finally, the result is restored to the original channel dimension and residually connected with the input to enhance feature representation capability.
[0107] Specifically, MSLA is placed after the upsampling stage of the generator (the first upsampling module and the second upsampling module), corresponding to the process of gradually recovering the high-resolution image from the low resolution. It is mainly used to capture long-range dependencies in the input feature map. By modeling the spatial information of the feature map, MSLA enables the features of each pixel to not only depend on local neighborhood information, but also capture global contextual information, thereby improving the expressive power of the model.
[0108] Specifically, for the MSLA module, assuming the input feature map is... , Where B represents the batch size, C represents the number of channels, and H×W represents the image spatial resolution.
[0109] Input Projecting the Q, K, and V values required for multi-head attention:
[0110]
[0111]
[0112]
[0113] Where h represents the number of attention heads. , is the number of channels per head. It's the compression ratio. This indicates that the average pooling operation extracts multi-scale features after downsampling. In this embodiment, the pooling kernel size is set to 2×2, and the stride is 2, therefore the output resolution is... .
[0114] In attention mechanisms, These are the query weight, key weight, and value weight, respectively. In the multi-scale linear attention structure of this embodiment, These are linear projection matrices that map input features to the Q, K, V spaces. These projection matrices are learnable parameters during training and determine the key capabilities of information extraction and reconstruction in the attention mechanism.
[0115] The computational complexity of traditional attention is ,in The calculation formula is as follows:
[0116]
[0117] The computational complexity of this model is... This model and By input The downsampled features obtained after spatial pooling are then transformed linearly, thus greatly reducing the computational complexity.
[0118] First of all Perform softmax normalization in both the channel and spatial dimensions:
[0119]
[0120] Next, calculate the attention output:
[0121]
[0122] In this model, softmax is directly applied to... Normalization is performed to avoid large-scale N×N similarity calculations, thereby reducing computational complexity. At the same time, downsampling is used to achieve multi-scale modeling, balancing global receptive field and computational efficiency.
[0123] Specifically, the MSLA processing used in the first upsampling module mainly processes low-level features to enhance the recovery capability of coarse structures.
[0124] Specifically, the MSLA used in the second upsampling module mainly processes fine features to help restore details and accurate stroke information.
[0125] Example 4:
[0126] Figure 8 This is a schematic diagram of the repair discriminator structure according to an embodiment of the present invention.
[0127] Specifically, the repair discriminator in this embodiment of the invention is a typical DCGAN discriminator, which includes multiple convolutional layers and activation functions.
[0128] Specifically, the image of the predicted intact Qin bamboo slip text generated by the repair generator will be repaired. (or intact training images) The input is fed into the repair discriminator, which gradually extracts features from the image through four convolutional layers (downsampling layers) and reduces the spatial resolution of the image. Finally, the convolutional layer outputs a single value to give an evaluation of the image's authenticity. The Sigmoid activation function is then used to map the output to the range [0, 1] to determine the true probability of the image.
[0129] Figure 9 This is a schematic diagram of the downsample layer structure according to an embodiment of the present invention.
[0130] A 4×4 convolutional kernel with stride=2 is used to achieve progressive downsampling, halving the resolution at each layer and increasing the number of channels from 64 to 512. Each convolutional kernel is activated by LeakyReLU after processing. The 5th layer uses a 3×3 convolution to compress the feature map into a 1×1 output, which is then activated by Sigmoid to generate the discriminant probability. All convolutional layers are trained using spectral normalization.
[0131] Throughout the training process, the repair generator and the repair discriminator are optimized alternately, and the model's parameters are optimized through training. The training process can be summarized as an adversarial game between the repair generator and the repair discriminator: the repair discriminator tries its best to distinguish between real and generated images, while the repair generator tries its best to generate real images, thereby "deceiving" the discriminator.
[0132] The loss of the repair discriminator consists of two parts: the binary classification loss of the intact training images. Binary classification loss for predicting intact Qin bamboo slip text images .
[0133]
[0134]
[0135] in, and Both use binary classification with cross-entropy loss.
[0136] The total loss from the repair assessment is:
[0137]
[0138] The goal of the repair generator is to minimize the error of the repair discriminator in judging the generated image, so that the generated image is more like a real image.
[0139] The loss function of the repair generator includes reconstruction loss. Combating losses and perceived loss .
[0140]
[0141] Reconstruction losses That is, generate an image With target image The mean square error loss between them.
[0142]
[0143] To generate adversarial loss, i.e., the generator aims to maximize... This makes its output close to 1.
[0144]
[0145] The definition of perceived loss is as follows: , , These are the feature maps at the th The number of channels, height, and width of a layer can be compared by selecting features from only a single layer; the perceptual loss can then be simplified to...
[0146]
[0147] The total loss of the generator is:
[0148]
[0149] in, and Represents the weighting factor, set , .
[0150] It should be noted that in the entire conditional repair generative network, the basic component structures such as convolutional layers are all existing technologies. The basic component structures used in different modules are the same, and their actual function is determined by the position of the basic components in the model.
[0151] Specifically, in each training step, the repair discriminator is trained using a predicted intact Qin bamboo slip text image and an intact training image. The repair discriminator is optimized by maximizing its prediction probability for the intact training image and minimizing its prediction probability for the predicted intact Qin bamboo slip text image. The repair generator is optimized based on its loss. The repair generator is trained by minimizing the difference between its predicted intact Qin bamboo slip text image and the intact training image and by trying its best to make the repair discriminator think that its output predicted intact Qin bamboo slip text image is the intact training image.
[0152] Specifically, after each iteration, the weights of the repair generator and the repair discriminator are updated through the backpropagation algorithm, and the optimizer is Adam.
[0153] In practical applications, the damaged Qin bamboo slip text image data in the damaged Qin bamboo slip text image dataset can be divided into training data and testing data. After the repair generator and repair discriminator are trained using the training data, the repair generator can be tested using the testing data. Based on the output results, PSNR and SSIM are calculated to evaluate the repair performance of the repair generator in this embodiment of the invention.
[0154] Example 5:
[0155] Figure 10 This is a schematic diagram comparing the physical image of Qin bamboo slip characters and the binarized image of Qin bamboo slip characters in an embodiment of the present invention.
[0156] Specifically, multiple images of Qin bamboo slip characters are extracted from data not used in training and testing. To meet the model data format requirements of this invention (input must be a binary image), the images can be processed as follows: the grayscale background of the Qin bamboo slip characters is removed and converted into Qin bamboo slip characters binary images. Specifically, the background removal operation first converts the input image into a grayscale image, reads the grayscale image, converts it into a single-channel image, and then performs a binarization operation, setting a threshold. (Can be set to 128 depending on actual needs), pixel value greater than or equal to Set the pixel value to 255, which is less than The pixel value is set to 0, thereby achieving the purpose of background removal.
[0157] Refer to the attached diagram. Figure 10, after binarization, the corresponding Qin bamboo slip character entity image forms a corresponding binary image of Qin bamboo slip characters. It should be noted that for the sake of clarity of illustration, Figure 10 The sampled Qin bamboo slip character entity image is a Qin bamboo slip character entity image with a relatively complete physical structure.
[0158] Figure 11 This is a schematic diagram of the repair example of the repair generator in the embodiment of the present invention. Among them, the first column is the unknown damaged Qin bamboo slip characters, the second column is the repaired Qin bamboo slip character image, and the third column is the manually repaired Qin bamboo slip character image (after binarization). Refer to the attached drawing Figure 11 It is shown that for different configurations of the character "ke" and different damage situations, the repaired Qin bamboo slip character image generated by the real-time repair generator of the present invention has a high similarity with the manually repaired Qin bamboo slip character image, and can well replace manual labor for automated Qin bamboo slip character repair work.
[0159] In summary, the present invention provides a method for generating a forged image of damaged Qin bamboo slip characters and an image repair method. The method for generating a forged image of damaged Qin bamboo slip characters uses all types of image data in the same period as the reference object for extracting the damage mask, extracts the damage mask through a generative adversarial damage network, and then processes the intact Qin bamboo slip character image with the extracted damage mask to obtain a highly realistic damaged Qin bamboo slip character image. This method provides a method that can effectively expand the number of damaged ancient character samples and has good practicability in machine learning applications in related fields.
[0160] The above has introduced in detail a method for generating a forged image of damaged Qin bamboo slip characters and an image repair method provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for generating images of damaged Qin Dynasty bamboo slips, characterized in that, include A damaged image dataset is constructed, which includes several damaged training data sets. Each damaged training data set includes a damaged training image and a corresponding intact training image. The intact training images include intact training images of Qin bamboo slips based on Qin bamboo slips and intact training images of non-Qin bamboo slips based on images of other works from the same period as Qin bamboo slips. The damaged training images include damaged training images of Qin bamboo slips based on Qin bamboo slips and damaged training images of non-Qin bamboo slips based on images of other works from the same period as Qin bamboo slips. An initialization of a generative adversarial network is performed. The generative adversarial network includes a mask generator, a damage generator, and a damage discriminator. The mask generator is used to output a predicted damage mask based on the input intact image. The damage generator is used to process the corresponding intact image based on the predicted damage mask to obtain a predicted damaged image. The damage discriminator is used to evaluate the authenticity of the input damaged image based on a preset damage evaluation index. The generative adversarial network is trained by iteratively training the network using damaged training data from the damaged image dataset until training is complete. Construct a mask dataset, and sequentially import the intact training images from the damaged image dataset into the trained generative adversarial destruction network. The mask generator in the generative adversarial destruction network sequentially generates corresponding predicted destruction masks, and all predicted destruction masks generated by the mask generator are stored in the mask dataset. The process involves generating a simulated image of Qin bamboo slip characters by randomly extracting a predicted damage mask from the mask dataset and processing any intact Qin bamboo slip character training image from the damaged image dataset to obtain a corresponding simulated damaged training image. This simulated damaged training image is the desired simulated image of Qin bamboo slip characters.
2. The method for generating images of damaged Qin bamboo slips as described in claim 1, characterized in that, The damaged training images have a preset fixed resolution.
3. The method for generating images of damaged Qin bamboo slips as described in claim 1, characterized in that, Both the damaged training image and the intact training image are binarized images.
4. A method for restoring damaged Qin Dynasty bamboo slip text images, comprising a training process and a restoration process, characterized in that, The training process includes: A dataset of real images of damaged Qin bamboo slips is constructed. The dataset of real images of damaged Qin bamboo slips includes several real images of damaged Qin bamboo slips. Each real image of damaged Qin bamboo slips includes a training real image of the damaged Qin bamboo slips and a corresponding intact training image. A dataset of images of Qin bamboo slips with inscriptions is constructed. The dataset includes several sets of images of Qin bamboo slips with inscriptions. Each set of images includes an image of Qin bamboo slips with inscriptions generated by the method of generating images of Qin bamboo slips with inscriptions according to any one of claims 1 to 3, and a corresponding intact training image. A dataset of images of damaged Qin bamboo slips is constructed by merging the dataset of real images of damaged Qin bamboo slips with the dataset of simulated images of damaged Qin bamboo slips. The dataset of images of damaged Qin bamboo slips includes several sets of images of damaged Qin bamboo slips. Each set of images includes a training image of a damaged Qin bamboo slip and a corresponding intact training image. The types of training images of damaged Qin bamboo slips include real training images of damaged Qin bamboo slips and simulated training images of damaged Qin bamboo slips. An initialization generative adversarial repair network is performed, which includes a repair generator and a repair discriminator. The repair generator is used to output a predicted image of intact Qin bamboo slips based on the input image of damaged Qin bamboo slips. The repair discriminator is used to evaluate the authenticity of the image of intact Qin bamboo slips input to the repair discriminator. The generative adversarial repair network is trained by iteratively training the network using training data from the damaged Qin bamboo slips text dataset until training is complete.
5. The method for restoring damaged Qin bamboo slip text images as described in claim 4, characterized in that, The repair process includes: An unknown damaged Qin bamboo slip character is input into the generative adversarial repair network. The generator of the generative adversarial repair network generates a predicted intact Qin bamboo slip character image, which is the repaired Qin bamboo slip character image after the unknown damaged Qin bamboo slip character has been repaired.
6. The method for restoring damaged Qin bamboo slip text images as described in claim 4, characterized in that, The repair generator comprises a first residual block, a second residual block, a bottleneck convolutional layer, a first upsampling module, a second upsampling module, a third upsampling module, a bilinear interpolation layer, and a normal convolutional layer, linked in sequence.
7. The method for restoring damaged Qin bamboo slip text images as described in claim 6, characterized in that, The first upsampling module and the second upsampling module are each equipped with a multi-scale linear attention module.
8. The method for restoring damaged Qin bamboo slip text images as described in claim 6 or 7, characterized in that, The loss function of the repair generator includes reconstruction loss, adversarial loss, and perceptual loss.
9. The method for restoring damaged Qin bamboo slip text images as described in claim 4, characterized in that, The repair discriminator is a DCGAN discriminator.
10. The method for restoring damaged Qin bamboo slip text images as described in claim 7, characterized in that, The loss function of the repair discriminator includes the binary classification loss of the predicted intact Qin bamboo slip text image and the binary classification loss of the intact training image.
Citation Information
Patent Citations
Qinxi character restoration method based on conditional generative adversarial network
CN116681604A
Hyperspectral ancient painting detection and recognition method based on deep learning
CN111291675A
Construction method and device of classification model for evaluating implant stability based on CBCT image data
CN115512167A