Ceramic cultural relic image highlight removing method based on Transform and diffusion model
Through the method based on Transformer and diffusion model, low-frequency and high-frequency information of ancient ceramic images are extracted, high-quality prompts are generated, and precise highlight removal and image repair are performed, which solves the problem of unsatisfactory de-highlight effect in the existing technology, and achieves efficient highlight removal and retention of detailed information.
Patent Information
- Application Number
- CN202510282779.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
When removing the highlights of ancient ceramics, the existing deep learning-based highlights have errors in restoring color and texture details, poor adaptability, and ignore the key low-frequency and high-frequency differences in the image, resulting in unsatisfactory de-highlighting effects.
The ceramic cultural relics image highlight removal method based on Transformer and diffusion model is adopted. The low-frequency and high-frequency information are extracted through the frequency prompt encoder, and high-quality prompts are generated in combination with the diffusion model. The Transformer highlight removal module is used for precise de-highlight processing, and the image repair is carried out through the prompt guidance of the reconstruction module.
Effectively removes the highlights of ancient ceramics, retains the detailed information of the image, improves the removal of highlights, avoids the appearance of visual defects, and has good applicability.
Smart Images

Figure CN120219259A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to image processing, in particular to a method for removing highlights from a ceramic cultural relic image based on a Transformer and a diffusion model. Background Art
[0002] As an important cultural heritage, ancient ceramics carry rich historical, cultural, artistic and scientific information, and are key physical materials for studying the social, economic and cultural development of ancient times. However, due to the erosion of time and the preservation environment, ancient ceramics face many risks of disease. Digital technology is needed to comprehensively and accurately record and protect them so that they can be better inherited and studied.
[0003] Surface details are the key basis for identifying the age, kiln, and production process of ancient ceramics. For example, the subtle differences in the body and glaze color of the five famous kiln porcelains in the Song Dynasty are crucial to judging the authenticity and kiln ownership. However, details such as the glaze, texture, and shape on the surface of ceramics are often strongly reflected by light, forming highlight areas with a full sense of gloss and almost no details of bright white or light-colored spots, which causes some fine painting patterns on ancient ceramics (for example, landscape and figure patterns on blue and white porcelain, floral patterns on pastel porcelain, etc.) to be covered up, and the delicacy of the lines and the layering of colors are weakened, which is very unfavorable for the identification and protection of cultural relics. Therefore, it is particularly important to remove the highlights of ancient ceramic cultural relics.
[0004] Early methods for removing image highlights mostly acquired multiple images by moving light sources or cameras, and then removed highlights through polarization filtering or color information. For example, Boult and Wolff used polarization filters to separate the reflection component from the grayscale image. However, due to hardware limitations and the inability to acquire multiple images at the same time, the adaptability was poor. Single-image highlight removal methods based on two-color reflection models, color space, and adjacent pixels also have many problems. For example, Shen's method will blur large-area detail information, and Akashi's method cannot completely remove highlights and will increase noise, resulting in missing detail information. It can be seen that the above traditional highlight removal methods cannot effectively solve the problem of hue and saturation ambiguity, and the highlight removal effect is not ideal for ancient ceramic images with complex texture details and rich colors.
[0005] In the deep learning convolutional neural network (CNN) model, it can automatically learn the feature representation of images from a large amount of data without manual feature extraction, and can capture complex patterns and abstract features in the images. This is very beneficial for processing rich texture, color and other detailed information in ancient ceramic images, and can better understand and remove the highlight area. There are already some models for removing highlights based on deep learning methods. For example, the unified framework for joint highlight detection and removal proposed by Fu et al. designed multiple extended spatial context feature aggregation (dscfa) modules to obtain context features at different scales, which can accurately detect the highlight position and remove it. However, there is still a difference between the detailed information and the ground truth map. In particular, when dealing with highlights in a colored lighting scene, the effect is still not ideal.
[0006] All in all, the existing deep learning-based highlight removal methods still have the following problems in the process of removing highlights from ancient ceramics: First, there are errors in restoring the original color and texture details of ancient ceramics and cannot be completely accurately restored, resulting in damage to the visual effect and information integrity of the image; Second, when facing highlight images of ancient ceramics with specific material, shape and texture characteristics, the adaptability is poor and cannot be well adjusted and optimized according to the characteristics of different ancient ceramics, thus affecting the highlight removal effect; Finally, the existing single-image highlight removal methods based on deep learning often ignore the key low-frequency and high-frequency differences in the image, thus affecting the effectiveness of its highlight removal, and the semantic disambiguation effect is not good, and visual defects such as black color blocks may appear in the output result. Summary of the Invention
[0007] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a highlight removal method for ancient ceramic cultural relic images based on Transformer and diffusion models, which can effectively remove highlights and retain the detailed information of the images, and has good applicability.
[0008] In order to achieve the above purpose, the present invention adopts the following technical solutions to implement:
[0009] A highlight removal method for ancient ceramic cultural relic images based on Transformer and diffusion models includes the following steps:
[0010] Step 1, preprocess the collected ancient ceramic highlight images to obtain the preprocessed ancient ceramic highlight images;
[0011] Step 2, after splicing the preprocessed ancient ceramic highlight images and the real ancient ceramic images without highlights, input them into the frequency prompt encoder for training, and use the trained frequency prompt encoder to encode the real ancient ceramic images without highlights into high-frequency prompts and low-frequency prompts;
[0012] Step 3, pre-train the diffusion model prompt generator. The diffusion model prompt generator can generate a prompt P for highlight removal based on the input high-frequency prompt and low-frequency prompt;
[0013] Step 4, construct a Transformer highlight removal module, including a prompt interaction and injection module PIIM and a prompt feed-forward network. The prompt interaction and injection module PIIM includes a prompt interaction module and a prompt injection module. The prompt interaction module includes a multi-head self-attention module, where:
[0014] The prompt interaction module first uses an adaptive average pooling operation and two linear layers in sequence to transform the input features Figure X l-1 into a shape identical to that of the prompt P which is expressed as:
[0015]
[0016] In the formula, W Q , W K and W V represent the projection matrices of the query Q, the key K, and the value V respectively; K p represents the vector of the key K;
[0017] Then, through the multi-head self-attention module, cross-attention interaction is performed on the prompt P and in the spatial dimension, which is expressed as:
[0018]
[0019] In the formula, α represents an optional factor for adjusting the output result of the Softmax function; represents the transpose of K p ; Adap() represents the adaptive average pooling operation; P′ represents the self-refined prompt;
[0020] Then, the P′ is refined through two learnable fusion parameters γ and β to obtain the refined prompt P r , which is expressed as:
[0021]
[0022] In the formula, ⊙ is the element-wise multiplication, μ and σ are the mean and standard deviation of P′ respectively, and γ and β are generated by f γ (P) and f β (P) respectively. f γ (P) and f β (P) are implemented through a linear layer with one intermediate layer normalization and one rectified linear unit layer;
[0023] The aforementioned prompt injection module injects the refined prompt P r into the feature Figure X l-1 to obtain the adjusted feature Figure X ′ l-1 , which is expressed as:
[0024] X′ l-1 =W1P r ⊙X l-1 +W2P r
[0025] wherein, W1 and W2 represent linear layers;
[0026] The aforementioned prompt feed-forward network performs feature dimension transformation and non-linear activation on the adjusted feature Figure X ′ l-1 to obtain the feature map I';
[0027] Step 5: Jointly train the diffusion model prompt generator and the Transformer highlight removal module. Use the trained diffusion model prompt generator to generate the prompt P for highlight removal. Under the guidance of the prompt P, the trained Transformer highlight removal module removes the highlights from the input feature map;
[0028] Step 6: Use the prompt-guided reconstruction module to repair and reconstruct the feature map I' to obtain the highlight-removed ancient ceramic image.
[0029] Furthermore, the specific process of the aforementioned Step 1 is as follows:
[0030] Step 1.1: Enhance the collected ancient ceramic highlight images to obtain the enhanced ancient ceramic highlight images;
[0031] Step 1.2: Crop a 256×256 pixel area from the center of the enhanced ancient ceramic highlight images to obtain the preprocessed ancient ceramic highlight images.
[0032] Furthermore, the enhancement in the aforementioned Step 1.1 includes geometric transformation, color transformation, and noise addition, where:
[0033] The geometric transformation includes rotation, flipping, and scaling, and the specific process is as follows:
[0034] Step 1.1.1.1: Randomly rotate the ancient ceramic highlight images at an angle between -30° and 30° to change the position and angle of the highlights;
[0035] Step 1.1.1.2: Perform horizontal flipping and vertical flipping operations on the ancient ceramic highlight images;
[0036] Step 1.1.1.3: Scale the high-gloss image of ancient ceramics according to the ratios of 0.8 and 1.2 to change the size and relative position of the high gloss.
[0037] The color transformation includes brightness, contrast, and hue adjustment. The specific process is as follows:
[0038] Step 1.1.2.1: Multiply the high-gloss image of ancient ceramics by a random brightness coefficient between 0.8 and 1.2 to adjust the overall brightness of the image.
[0039] Step 1.1.2.2: Change the contrast of the high-gloss image of ancient ceramics to highlight or weaken the high-gloss area.
[0040] Step 1.1.2.3: Adjust the hue value of the high-gloss image of ancient ceramics in the HSV color space.
[0041] The noise addition is to add Gaussian noise with a mean of 0 and a variance of 0.01 to the high-gloss image of ancient ceramics.
[0042] Furthermore, the frequency cue encoder in step 2 includes wavelet transform and a dual-branch encoder. Its processing process for the input image is as follows:
[0043] Step 2.1: The wavelet transform converts the input image from the spatial domain to the frequency domain and decomposes it into high-frequency cues and low-frequency cues.
[0044] Step 2.2: The dual-branch encoder encodes the high-frequency cues and low-frequency cues to extract and compress the features of the corresponding frequency information.
[0045] Furthermore, the processing process of the diffusion model cue generator in step 3 for the input high-frequency cues and low-frequency cues is as follows:
[0046] Step 3.1: Cue diffusion
[0047] Step 3.1.1: First, continuously add Gaussian noise to the high-frequency cues and low-frequency cues through the denoising network, convert the high-frequency cues and low-frequency cues into standard Gaussian noise ∈, and then add the high-frequency cues and low-frequency cues at time t to the standard Gaussian noise ∈ respectively to obtain the noisy high-frequency cues and noisy low-frequency cues which are respectively expressed as:
[0048]
[0049] Step 3.1.2: Take and The input denoising network ∈ θ , to obtain the predicted noise, and then use the high-frequency diffusion loss of the l-th layer to constrain the difference between the standard Gaussian noise ∈ and the predicted noise, expressed as:
[0050]
[0051] Meanwhile, take and the input denoising network ∈ θ , to obtain the predicted noise, and use the low-frequency diffusion loss of the l-th layer to constrain the difference between the standard Gaussian noise ∈ and the predicted noise, expressed as:
[0052]
[0053] Step 3.2, Prompt generation
[0054] First, input the denoised low-frequency prompt and the low-frequency prompt into the denoising network ∈ θ , to obtain the predicted noise, and then subtract the predicted noise from the denoised prompt at time t to obtain the denoised low-frequency prompt at the next time t+1 expressed as:
[0055]
[0056] Meanwhile, first input the denoised high-frequency prompt and into the denoising network ∈ θ , to obtain the predicted noise, and then subtract the predicted noise from the denoised high-frequency prompt at time t to obtain the denoised high-frequency prompt at the next time t+1 expressed as:
[0057]
[0058] Furthermore, the joint training loss adopted for jointly training the diffusion model prompt generator and the Transformer highlight removal module in step 5 is expressed as:
[0059]
[0060] In the formula, is the low-frequency loss, is the high-frequency loss, is the per-pixel loss generated by the Transformer highlight removal module, expressed as:
[0061]
[0062] In the formula, I gt represents the real ancient ceramic image without highlights, and I' is the feature map output by the Transformer highlight removal module.
[0063] Furthermore, the specific process of step 6 is as follows: First, use the multi-layer perception mechanism MLP to mine the relationship between the features of the feature map I', then use transposed convolution for upsampling to increase the size of the feature map I', and then perform two-layer 3×3 convolution processing in sequence. After each convolution processing, use the ReLU activation function for non-linear transformation, and then perform a 3×3 convolution processing and map the feature values through the Tanh activation function to obtain the ancient ceramic image without highlights.
[0064] Compared with the prior art, the present invention has the following technical effects:
[0065] First, adopt a decoupled network model framework, which is divided into two stages: prompt generation and prompt-guided restoration as a whole, and is more targeted for the highlight processing of ancient ceramic images; second, use the low-frequency information and high-frequency information output by the frequency prompt encoder as prompts, and the Transformer highlight removal module can effectively distinguish the highlight and background features; third, use the diffusion model to process the low-frequency information and high-frequency information to generate high-quality prompts, providing reliable guidance for removing highlights and improving the highlight removal effect; fourth, introduce the prompt interaction and injection module PIIM into the Transformer prompt block, and use its self-attention mechanism to grasp the long-range dependence, focusing more on the image regions and features related to highlight removal, comprehensively removing highlights and retaining the original material, color, texture and other detail information of the image; in summary, the present invention continuously refines the prompts output by the frequency prompt encoder through the diffusion model prompt generator, uses the refined prompts to guide the Transformer highlight removal module to accurately remove the highlights in the ancient ceramic highlight image and completely retain the detail information, and then outputs the complete and clear ancient ceramic image without highlights through the prompt-guided reconstruction module, improving the highlight removal effect and preventing the appearance of visual defects. Description of the Drawings
[0066] Figure 1 is a schematic structural diagram of the overall network model of the present invention;
[0067] Figure 2 is the selected ancient ceramic highlight image of the present invention;
[0068] Figure 3 is the framework diagram of the overall network model of the present invention;
[0069] Figure 4 is the framework diagram of the diffusion model prompt generator of the present invention;
[0070] Figure 5 It is a schematic structural diagram of the Prompt Interaction and Injection Module (PIIM) of the present invention;
[0071] Figure 6 It is a schematic structural diagram of the multi-head self-attention module of the present invention;
[0072] Figure 7 It is a schematic structural diagram of the prompt feed-forward network of the present invention;
[0073] Figure 8 It is a schematic structural diagram of the prompt-guided reconstruction module of the present invention;
[0074] Figure 9 The real ancient ceramic images without highlights selected by the present invention;
[0075] Figure 10 They are the original ceramic highlight image A and the predicted highlight-free ceramic image B selected in the embodiment of the present invention;
[0076] Figure 11 They are the original ceramic highlight image C and the predicted highlight-free ceramic image D selected in the embodiment of the present invention. Detailed implementation manners
[0077] The following further elaborates on the specific content of the present invention in conjunction with embodiments.
[0078] This embodiment provides a method for removing highlights from ceramic cultural relic images based on Transformer and diffusion models. The network framework used is as Figure 1 and Figure 3 shown, including a frequency prompt encoder and a prompt transformer. Among them, the prompt transformer includes a diffusion model prompt generator, a Transformer highlight removal module, and a prompt-guided reconstruction module.
[0079] As Figure 1 shown, a method for removing highlights from ceramic cultural relic images based on Transformer and diffusion models includes the following steps:
[0080] Step 1: Preprocess the collected ancient ceramic highlight images to obtain preprocessed ancient ceramic highlight images. The specific process is as follows:
[0081] Step 1.1: Enhance the collected ancient ceramic highlight images to obtain the enhanced ancient ceramic highlight images as Figure 2 shown;
[0082] The enhancement includes geometric transformation, color transformation, and noise addition, where:
[0083] The geometric transformation includes rotation, flipping, and scaling. The specific process is as follows:
[0084] Step 1.1.1.1: Rotate the high-gloss image of ancient ceramics randomly at an angle between -30° and 30°, changing the position and angle of the high-gloss, enabling the model to adapt to high-gloss in different directions, thereby enhancing the generalization ability of the model and preventing the model from relying on a fixed pattern of high-gloss position.
[0085] Step 1.1.1.2: Perform horizontal flipping and vertical flipping operations on the high-gloss image of ancient ceramics. For high-gloss generated by symmetric objects, the flipping operation can simulate the high-gloss situation of the object in different directions, thereby effectively increasing the data volume.
[0086] Step 1.1.1.3: Scale the high-gloss image of ancient ceramics at ratios of 0.8 and 1.2, changing the size and relative position of the high-gloss, enabling the model to learn high-gloss features at different scales, which helps the model to effectively process images with different resolutions or high-gloss regions of different sizes.
[0087] The color transformation includes brightness, contrast, and hue adjustment. The specific process is as follows:
[0088] Step 1.1.2.1: Multiply the high-gloss image of ancient ceramics by a random brightness coefficient between 0.8 and 1.2 to adjust the overall brightness of the image and change the contrast between the high-gloss and the surrounding area. This can simulate the high-gloss situation under different light intensities, enabling the model to adapt to various lighting environments.
[0089] Step 1.1.2.2: Use contrast stretching to change the contrast of the high-gloss image of ancient ceramics, highlighting or weakening the high-gloss area, helping the model learn high-gloss features under different contrasts, and enhancing the robustness of the model to high-gloss.
[0090] Step 1.1.2.3: Adjust the hue value of the high-gloss image of ancient ceramics in the HSV color space, enabling the model to process high-gloss with different color tendencies.
[0091] The noise addition is to add Gaussian noise with a mean of 0 and a variance of 0.01 to the high-gloss image of ancient ceramics, simulating the noise interference in the actual shooting environment, enabling the model to accurately remove the high-gloss even in the presence of noise.
[0092] Step 1.2: Crop a 256×256 pixel area from the center of the enhanced high-gloss image of ancient ceramics to obtain the preprocessed high-gloss image of ancient ceramics.
[0093] Step 2: The preprocessed high-gloss image of ancient ceramics as shown in Figure 2 and as shown in Figure 9After the spliced real ancient ceramic images without highlights are input into the frequency prompt encoder for training, the frequency prompt encoder can learn the frequency feature information related to highlight removal in the images. The trained frequency prompt encoder encodes the real ancient ceramic images without highlights into high-frequency prompts and low-frequency prompts;
[0094] The frequency prompt encoder includes wavelet transform and a dual-branch encoder. Its processing process for the input image is as follows:
[0095] Step 2.1: First, the wavelet transform converts the input image from the spatial domain to the frequency domain, enabling better capture of the low-frequency and high-frequency information in the image. Then, the image is decomposed into sub-bands of different frequencies to obtain high-frequency prompts and low-frequency prompts, providing frequency feature prompts for subsequent encoding;
[0096] Step 2.2: The dual-branch encoder encodes the high-frequency prompts and low-frequency prompts. Each branch of the dual-branch encoder includes a series of convolutional layers, pooling layers, and activation functions for extracting and compressing the features of the corresponding frequency information. The dual-branch structure can respectively focus on low-frequency features and high-frequency features, thus effectively extracting the frequency prompt information in the image;
[0097] Step 3: Use an existing diffusion model as a prompt generator, called the diffusion model prompt generator. Its structure is as Figure 4 shown. Pretrain the diffusion model prompt generator. The trained diffusion model prompt generator can generate a prompt P for highlight removal based on the input high-frequency prompts and low-frequency prompts, providing guiding information for removing the highlights in the image;
[0098] The processing process of the diffusion model prompt generator for the input high-frequency prompts and low-frequency prompts is as follows:
[0099] Step 3.1: Prompt diffusion
[0100] Step 3.1.1: First, continuously add Gaussian noise to the high-frequency prompt and the low-frequency prompt through the denoising network. The features in the high-frequency prompt and the low-frequency prompt gradually disappear and are converted into standard Gaussian noise ∈. Then, add the high-frequency prompt and the low-frequency prompt at time t to the standard Gaussian noise ∈ respectively to obtain the noisy high-frequency prompt and the noisy low-frequency prompt at the next time t + 1, which are respectively expressed as:
[0101]
[0102] The high-frequency prompt and low-frequency prompts are condition prompts contaminated by high light for the output of the frequency encoder;
[0103] Step 3.1.2, Take and input into the denoising network ∈ θ , obtain the predicted noise, and then use the high-frequency diffusion loss of the l-th layer to constrain the difference between the standard Gaussian noise ∈ and the predicted noise, expressed as:
[0104]
[0105] Meanwhile, take and input into the denoising network ∈ θ , obtain the predicted noise, use the low-frequency diffusion loss of the l-th layer to constrain the difference between the standard Gaussian noise ∈ and the predicted noise, expressed as:
[0106]
[0107] Step 3.2, Prompt generation
[0108] First, take the denoised low-frequency prompt and the low-frequency prompt input into the denoising network ∈ θ , obtain the predicted noise, and then subtract the predicted noise from the denoised prompt at time t to obtain the denoised low-frequency prompt at the next time t+1 expressed as:
[0109]
[0110] Meanwhile, first take the denoised high-frequency prompt and input into the denoising network ∈ θ , obtain the predicted noise, and then subtract the predicted noise from the denoised high-frequency prompt at time t to obtain the denoised high-frequency prompt at the next time t+1 expressed as:
[0111]
[0112] Step 4, Construct the Transformer high-light removal module, including the prompt interaction and injection module PIIM as shown in Figure 5 and the Figure 7The shown prompt feedforward network optimizes the input features using high-frequency and low-frequency prompts before generating the query Q, key K, and value V by introducing prompt interaction and injection module PIIM in the Transformer, as Figure 5 shown, the prompt interaction and injection module PIIM includes a prompt interaction module and a prompt injection module, and the prompt interaction module includes Figure 6 the multi-head self-attention module shown, where:
[0113] The prompt interaction module first uses adaptive average pooling operation and two-layer linear layer in sequence to convert the input features Figure X l-1 into the same shape as the prompt P, which is expressed as:
[0114]
[0115] In the formula, W Q , W K and W V represent the projection matrices of query Q, key K, and value V respectively; K p represents the vector of key K;
[0116] Then, through the multi-head self-attention module shown, cross-attention interaction is performed on the prompt P and Figure 6 in the spatial dimension to guide the multi-head self-attention module to focus more on the image regions and features related to highlight removal, thereby enhancing the highlight removal effect, which is expressed as:
[0117]
[0118] In the formula, α represents an optional factor for adjusting the output result of the Softmax function; represents the transpose of K p ; Adap() represents the adaptive average pooling operation; P′ represents the self-refined prompt;
[0119] Then, the P′ is refined through two learnable fusion parameters γ and β to obtain the refined prompt P r , which is expressed as:
[0120]
[0121] In the formula, ⊙ is the element-wise multiplication, μ and σ are the mean and standard deviation of P′ respectively, and γ and β are generated by f γ (P) and f β (P) respectively, and f γ (P) and f β (P) is implemented through a linear layer with intermediate layer normalization and a LeakyReLU layer, enabling the model to prioritize the highlight regions and the surrounding background information when processing images, thereby more accurately removing the highlights and restoring a clear background image;
[0122] The prompt injection module injects the refined prompt P r into the features Figure X l-1 to obtain the adjusted features Figure X ′ l-1 , which is expressed as:
[0123] X′ l-1 =W1P r ⊙X l-1 +W2P r
[0124] In the formula, W1 and W2 represent linear layers;
[0125] As Figure 7 shown, an existing feed-forward network PFFN is used as the prompt feed-forward network, and the prompt feed-forward network performs feature dimension transformation and non-linear activation on the adjusted features Figure X ′ l-1 to obtain the feature map I';
[0126] The prompt feed-forward network outputs a feature map I' with the same feature dimension as the adjusted features Figure X ′ l-1 to facilitate the subsequent image reconstruction operation smoothly;
[0127] Step 5: Jointly train the diffusion model prompt generator and the Transformer highlight removal module. Use the trained diffusion model prompt generator to generate a prompt P for highlight removal, and the trained Transformer highlight removal module removes the highlights from the feature map under the guidance of the prompt P;
[0128] The joint training loss used for jointly training the diffusion model prompt generator and the Transformer highlight removal module is expressed as:
[0129]
[0130] In the formula, is the low-frequency loss, is the high-frequency loss, is the per-pixel loss generated by the Transformer highlight removal module, which is expressed as:
[0131]
[0132] In the formula, I gt represents a real ancient ceramic image without highlights, and I' is the feature map output by the Transformer highlight removal module;
[0133] Step 6: Use Figure 8 The prompt-guided reconstruction module shown in the figure repairs and reconstructs the feature map I' to obtain a de-highlighted ancient ceramic image. The specific process is: first, the multi-layer perception mechanism MLP is used to explore the complex relationship between the features of the feature map I' to improve the model's ability to express image features, and then transposed convolution is used for upsampling to increase the size of the feature map I', so that the size of the processed low-resolution feature map I' is consistent with the size of the original ancient ceramic highlight image, which is convenient for restoring the complete image information. Then, the feature map is subjected to three-layer alkaline convolution processing. The specific process is: first, two layers of convolutional layers with a convolution kernel of 3×3, a step size of 1 and a padding of 1 are convolved in sequence. After each convolution, the ReLU activation function is used for nonlinear transformation, and then a layer of convolutional layer with a convolution kernel of 3×3, a step size of 1 and a padding of 1 is convolved, and the eigenvalues are mapped to an appropriate range through the Tanh activation function to obtain a de-highlighted ancient ceramic image.
[0134] In order to verify the effectiveness of the highlight removal method for ceramic artifact images based on Transformer and diffusion model proposed in this example, Figure 10 The original ceramic highlight image A in the input is as follows Figure 1 The network model shown in the figure has the following prediction results: Figure 10 From the non-highlight ceramic image B, it can be seen that the large area of highlight on image A has been removed, and no black information blocks or noise information are generated, which does not affect the overall appearance of the ceramic image. Figure 11 The original ceramic highlight image C in the input is as follows Figure 1 The network model shown in the figure has the following prediction results: Figure 11 From the highlight-free ceramic image D in the image C, it can be seen that the highlight area of the image C has been completely removed, and the information of the non-highlight area is completely retained; it can be seen that the highlight removal method for ceramic cultural relic images based on Transformer and diffusion model proposed in this embodiment can better restore the image's texture, color and other detail information during highlight removal, and the image does not have distortion, and has a good highlight removal effect.
Claims
1. A method for removing highlights from ceramic artifact images based on Transformer and diffusion model, characterized in that: The steps include: Step 1, preprocessing the collected ancient ceramic highlight images to obtain preprocessed ancient ceramic highlight images; Step 2: After the pre-processed ancient ceramic highlight image and the real ancient ceramic image without highlight are spliced, a frequency prompt encoder is input for training, and the trained frequency prompt encoder is used to encode the real ancient ceramic image without highlight into high-frequency prompts and low-frequency prompts; Step 3: pre-training a diffusion model prompt generator, where the diffusion model prompt generator can generate a prompt P for highlight removal based on the input high-frequency prompts and low-frequency prompts; Step 4: construct a Transformer highlight removal module, including a prompt interaction and injection module PIIM and a prompt feedforward network. The prompt interaction and injection module PIIM includes a prompt interaction module and a prompt injection module. The prompt interaction module includes a multi-head self-attention module, wherein: The prompt interaction module first uses an adaptive average pooling operation and two linear layers to transform the input feature map X l-1 Convert to the same shape as prompt P It is expressed as: K p =W K P Where W Q , W K and W V The projection matrices representing query Q, key K, and value V respectively; K p A vector representing the key K; Then, the multi-head self-attention module is used to analyze the prompts P and Perform cross attention interaction, expressed as: In the formula, α represents an optional factor used to adjust the output result of the Softmax function; K p The transpose of ; Adap( ) represents the adaptive average pooling operation; P′ represents the hint after self-refinement; Then, P′ is refined by two learnable fusion parameters γ and β to obtain the refined hint P r , expressed as: Where ⊙ is the element-by-element multiplication, μ and σ are the mean and standard deviation of P′, and γ and β are respectively given by f γ (P) and f β (P) generates, f γ (P) and f β (P) is implemented by a linear layer with an intermediate normalization layer and a rectified linear unit layer; The prompt injection module injects the refined prompt P r Inject feature map X l-1 , get the adjusted feature map X′ l-1 , expressed as: X′ l-1 =W1P r ⊙X l-1 +W2P r Where W1 and W2 represent linear layers; The prompt feed-forward network adjusts the feature map X′ l-1 Perform feature dimension transformation and nonlinear activation to obtain feature map I'; Step 5: Jointly train the diffusion model prompt generator and the Transformer highlight removal module. Use the trained diffusion model prompt generator to generate a prompt P for highlight removal. Under the guidance of the prompt P, the trained Transformer highlight removal module removes the highlights of the input feature map. Step 6: Use the prompt-guided reconstruction module to repair and reconstruct the feature map I' to obtain a de-highlighted ancient ceramic image.
2. The method for removing highlight from ceramic cultural relic images based on Transformer and diffusion model according to claim 1, characterized in that: The specific process of step 1 is as follows: Step 1.1, enhancing the collected ancient ceramic highlight image to obtain an enhanced ancient ceramic highlight image; Step 1.2: Crop a 256×256 pixel area from the center of the enhanced ancient ceramic highlight image to obtain the preprocessed ancient ceramic highlight image.
3. The method for removing highlight from ceramic cultural relic images based on Transformer and diffusion model according to claim 2, characterized in that: The enhancement of step 1.1 includes geometric transformation, color transformation and noise addition, wherein: The geometric transformation includes rotation, flipping and scaling, and the specific process is as follows: Step 1.1.1.1, rotate the highlight image of ancient ceramics at a random angle between -30° and 30° to change the position and angle of the highlight; Step 1.1.1.2, perform horizontal and vertical flip operations on the ancient ceramic highlight image; Step 1.1.1.3, scale the highlight image of ancient ceramics by 0.8 and 1.2 times to change the size and relative position of the highlight; The color transformation includes brightness, contrast and hue adjustment, and the specific process is as follows: Step 1.1.2.1, multiply the ancient ceramic highlight image by a random brightness coefficient of 0.8 to 1.2 to adjust the overall brightness of the image; Step 1.1.2.2, change the contrast of the highlight image of ancient ceramics to highlight or weaken the highlight area; Step 1.1.2.3, adjust the hue value of the ancient ceramic highlight image in the HSV color space; The noise addition is to add Gaussian noise with a mean of 0 and a variance of 0.01 to the ancient ceramic highlight image.
4. The method for removing highlight from ceramic cultural relic images based on Transformer and diffusion model according to claim 1, characterized in that: The frequency prompt encoder in step 2 includes wavelet transform and dual-branch encoder, and the processing process of the input image is as follows: Step 2.1, wavelet transform converts the input image from the spatial domain to the frequency domain and decomposes it into high-frequency cues and low-frequency cues; Step 2.2: The dual-branch encoder encodes high-frequency cues and low-frequency cues to extract and compress the features of the corresponding frequency information.
5. The method for removing highlight from ceramic cultural relic images based on Transformer and diffusion model according to claim 1, characterized in that: The diffusion model prompt generator in step 3 processes the input high-frequency prompts and low-frequency prompts as follows: Step 3.1: Prompt diffusion Step 3.1.1: First, continuously prompt the high frequency through the denoising network and low frequency prompts Add Gaussian noise to reduce the high frequency and low frequency prompts Convert to standard Gaussian noise ∈, and then convert the high-frequency prompt at time t and low frequency prompts Add them to the standard Gaussian noise ∈ to get the noisy high-frequency prompt at the next time t+1 And noisy low frequency prompt Respectively expressed as: Step 3.1.2: and Input denoising network∈ θ , get the predicted noise, and then use the high-frequency diffusion loss of the lth layer The difference between the standard Gaussian noise ∈ and the predicted noise is constrained, expressed as: At the same time, and Input denoising network∈ θ , get the predicted noise, using the low-frequency diffusion loss of the lth layer The difference between the standard Gaussian noise ∈ and the predicted noise is constrained, expressed as: Step 3.2: Prompt Generation First, remove the low frequency and low frequency prompts Input denoising network∈ θ , get the predicted noise, and then get the denoising hint from time t Subtract the predicted noise from the denoised low-frequency prompt at the next time t+1 It is expressed as: At the same time, first remove the high-frequency denoising prompt and Input denoising network∈ θ , get the predicted noise, and then denoise the high-frequency hint at time t Subtract the predicted noise from the denoised high-frequency prompt at the next time t+1 It is expressed as:
6. The method for removing highlights from ceramic cultural relic images based on Transformer and diffusion model according to claim 5, characterized in that: The joint training loss adopted in step 5 for jointly training the diffusion model hint generator and the Transformer highlight removal module is expressed as: In the formula, is the low frequency loss, is the high frequency loss, The pixel-by-pixel loss generated by the Transformer highlight removal module is expressed as: In the formula, I gt Represents a real ancient ceramic image without highlights, and I' is the feature map output by the Transformer highlight removal module.
7. The method for removing highlight from ceramic cultural relic images based on Transformer and diffusion model according to claim 1, characterized in that: The specific process of step 6 is: first use the multi-layer perception mechanism MLP to mine the relationship between the features of the feature map I', then use transposed convolution to upsample to increase the size of the feature map I', and then perform two layers of 3×3 convolution processing in sequence. After each convolution processing, the ReLU activation function is used for nonlinear transformation, and then a layer of 3×3 convolution processing is performed, and the eigenvalues are mapped through the Tanh activation function to obtain a de-highlighted ancient ceramic image.
Citation Information
Cited By
Image restoration method and system based on self-content-guided diffusion model
CN121685336A