Glass color migration method based on dynamic weight

By introducing a bi-branch feature extraction structure and a dynamic weighting function for visual simulation operations, the problem of illumination robustness in glass defect detection is solved, achieving image generation and content consistency of glass defects under different light source conditions, thus improving the model's adaptability and detection performance.

CN120912702APending Publication Date: 2025-11-07ANHUI LANSHI GLASS TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511021292.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In glass defect detection, existing technologies suffer from illumination robustness issues, leading to unstable image imaging and affecting the effectiveness and reliability of the detection model. Furthermore, existing methods struggle to maintain content consistency between the generated image and the content image under different light source conditions.

Method used

By employing a bi-branch feature extraction structure based on reversible networks and visual simulation operations, a dynamic weighting function is designed. Fine-grained and coarse-grained features are fused through a bi-branch dynamic weighting component to construct a color transfer model, thereby achieving image generation under different lighting conditions.

Benefits of technology

It effectively captures both fine-grained and coarse-grained features of images, solving the problem of difficulty in acquiring samples from different light sources for glass defect images, maintaining the consistency of content between the generated image and the content image, and improving the model's adaptability and detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912702A_ABST
    Figure CN120912702A_ABST
Patent Text Reader

Abstract

The invention discloses a glass color migration method based on dynamic weight, and the method comprises the steps: S1, building a data set: collecting data, building the data set, and carrying out the preprocessing of the data in the data set; s2, constructing and training a model: performing improvement based on a reversible network, and introducing a double-branch feature extraction structure and visual simulation operation to construct the model; training the model to obtain a trained glass color migration model based on the dynamic weight; and S3, outputting a result: inputting the style image and the content image into the glass color migration model trained in the step S2 and based on the dynamic weight at the same time, and generating an image of a target color. According to the method, a double-branch feature extraction structure and visual simulation operation are introduced, a dynamic weighting function is designed to fuse fine-grained features and coarse-grained features respectively, glass color migration is more effectively achieved through a double-branch dynamic weighting assembly, the problem that different light source samples are difficult to collect for glass defect images is solved, and the accuracy of glass defect detection is improved. And the method has higher application value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image style transfer, and particularly relates to a glass color transfer method based on dynamic weights. BACKGROUND

[0002] In the process of industrial automation, machine vision-based defect detection technology plays a core role. However, in the production line of high-reflective materials such as glass, the application of this technology faces a common problem of illumination robustness. The selection of light sources in the production environment, the change of illumination intensity and illumination conditions are the main sources of unstable imaging of glass surface images, which directly affects the color space of the image, and thus seriously weakens the effectiveness and reliability of existing defect detection models. In the face of this challenge, the mainstream approach in the current industry is to use a data-driven fine-tuning strategy. This method relies on the reconstruction of a dataset under new light conditions, which is not only a costly and time-consuming process, but frequent model training also introduces uncertainty, which may lead to fluctuations or even degradation of detection performance.

[0003] Currently, existing methods mainly use encoder-decoder architecture to realize the style transfer of specified content images and style images, such as AesPA-Net, MicroAST, and ArtFlow, etc. Due to the reflective and single color nature of glass pictures, when these models are used for glass datasets, the generated images may lose details and it is difficult to maintain the content consistency between the content image and the generated image. AesPA-Net proposes a pattern repeatability measurement method to quantify the degree of pattern repetition in the style image, and finds the best balance point between local and global style expression to improve the quality of the stylized image. However, it is difficult to accurately find this best point in practical applications, which may lead to loss of details.

[0004] MicroAST designs two micro-encoders to extract content and style features, and a micro-decoder to generate the final image. Limited by the limited expressive power of micro-networks, the synthesized images are prone to lose details. ArtFlow uses reversible neural flow to realize style transfer, which maps the content image and style image to the latent space through forward inference, and then uses WCT or AdaIN for style transfer in the backward inference process. Due to the possibility of introducing redundant information in the reversible flow structure, the generated images may lose details.

[0005] The Chinese patent document CN119810229A discloses a glass image color migration system, migration method, device and medium, belonging to the technical field of image style migration. The glass image color migration method comprises the following steps: step S100, acquiring an actual glass image; step S200, receiving the actual glass image, and performing color migration on the actual glass image by using a glass image color migration model to generate an actual glass color migration image. The technical solution mainly reduces the influence of color changes in the surrounding environment on the detection results of a subsequent glass quality image monitoring model, and improves the generalization ability of the glass quality image monitoring model, but does not solve the problem of difficulty in collecting different light source samples of glass defect images.

[0006] Therefore, given a target style image and a content image, how to maintain the content consistency of the generated image and the content image in style migration is an important challenge. It is necessary to develop a glass color migration method based on dynamic weights, which can generate glass defect images of corresponding colors according to target images and maintain the content consistency of the generated images and the content images, to solve the problem of difficulty in collecting different light source samples of glass defect images. SUMMARY

[0007] The purpose of the present application is to provide a glass color migration method based on dynamic weights, which can generate glass defect images of corresponding colors according to target images and maintain the content consistency of the generated images and the content images, to solve the problem of difficulty in collecting different light source samples of glass defect images.

[0008] To solve the above technical problems, the technical solution adopted by the present application is as follows: the glass color migration method based on dynamic weights comprises the following steps:

[0009] S1, constructing a data set: collecting data, constructing a data set, and preprocessing the data in the data set, adjusting the images to a uniform size, and dividing them into a training set and a test set;

[0010] S2, constructing and training a model: improving based on a reversible network, introducing a double-branch feature extraction structure and a visual simulation operation to construct a model; and training the model to obtain a trained glass color migration model based on dynamic weights;

[0011] S3, outputting results: inputting the style image and the content image into the glass color migration model based on dynamic weights trained in step S2 at the same time to generate an image of a target color.

[0012] By adopting the technical scheme, the double-branch feature extraction structure and visual simulation operation are introduced based on the reversible network, in the forward inference process, the fine-grained and coarse-grained features in the image can be effectively captured, and the double-branch dynamic weighting function is designed to fuse the fine-grained and coarse-grained features respectively, the color migration model is constructed, and the double-branch dynamic weighting component (from the double-branch feature extraction to the dynamic weighting function) is formed, and the target is to realize the cross-illumination condition image generation based on a single sample; by learning the color features and content features between different images, the model can migrate the color of any image to the specified content image, so as to generate a large amount of training data meeting the new illumination condition at a very low cost; in particular, the color migration can also be effectively realized for the industrial glass images with small data quantity and high data similarity; in this way, the data enhancement method is expected to solve the data acquisition problem at the source, and provide a solid guarantee for maintaining the long-term and stable high-performance of the detection model, so as to fundamentally improve the adaptability of the model.

[0013] Preferably, the specific step of the step S2 is:

[0014] S21: constructing a model, specifically:

[0015] S211: introducing a double-branch feature extraction structure composed of two convolution kernels with different sizes to extract fine-grained and coarse-grained features at the same time;

[0016] S212: then, visual simulation is performed through the up-sampling unit and the down-sampling unit respectively;

[0017] S213: introducing a channel attention unit to give weights to the two branches, and designing a dynamic weight function to fuse the extracted fine-grained and coarse-grained features respectively;

[0018] S214: finally, the fine-grained and coarse-grained features processed by the dynamic weight function are subjected to a cascade operation;

[0019] S22: training the model: inputting the training set for training, and iterating the model through the loss function to obtain the trained glass color migration model based on the dynamic weight, and then testing the model performance on the test set:

[0020] Preferably, the specific step of the step S211 is:

[0021] S211: in the forward inference process, the content image and the style image are inputted at the same time, and after the padding unit, they are divided into two parts along the channel dimension, one part is inputted into the projection network composed of 3 convolution layers to obtain the feature map;

[0022] S212: using a convolution kernel size of 3x3 and a convolution kernel size of 5x5 to process the feature map obtained in step S211, respectively obtaining the corresponding fine-grained feature map and coarse-grained feature map. That is, two branches are formed, one branch is to extract fine-grained feature map using a convolution kernel size of 3x3, and the other branch is to extract coarse-grained feature map using a convolution kernel size of 5x5.

[0023] Preferably, in step S212, the fine-grained feature map is simulated through the up-sampling unit to simulate the human "close look" visual process; the coarse-grained feature map is simulated through the down-sampling unit to simulate the human "far look" visual process; and the key features are emphasized; the specific steps are:

[0024] S2121: using interpolation and a convolution kernel size of 3x3 to extract the fine-grained feature map F f up-sampling and enlarging it by two times to simulate the human "close look" visual process, the formula is:

[0025] F fu = Conv(Interpolate(F f , scale_factor = 2));

[0026] Wherein, F fu is the fine-grained simulated visual feature map; Conv represents convolution; Interpolate represents interpolation up-sampling; scale_factor = 2 represents enlarging by 2 times;

[0027] S2122: using a convolution kernel size of 5x5 to extract the coarse-grained feature map F c down-sampling and reducing it by two times to simulate the human "far look" visual process, the formula is:

[0028] F cd = Conv(F c , stride = 2);

[0029] Wherein, F cd is the coarse-grained simulated visual feature map; Conv represents convolution; stride = 2 represents reducing by 2 times.

[0030] Preferably, the specific steps of step S213 are:

[0031] S2131: introducing a channel attention module, which is composed of a pooling layer, a convolution layer and an activation function;

[0032] S2132: the fine-grained feature map F f , the fine-grained simulated visual feature map F fu, the coarse-grained feature map F c and the coarse-grained simulated visual feature map F cd are respectively input into the channel attention module to calculate the channel attention weight, and four corresponding channel attention weights are obtained.

[0033] The fine-grained feature map F f and the fine-grained simulated visual feature map F fu are calculated, and the channel attention weights are added to obtain the final channel attention weight W f of the fine-grained feature map.

[0034] The coarse-grained feature map F c and the coarse-grained simulated visual feature map F cd are calculated, and the channel attention weights are added to obtain the final channel attention weight W c of the coarse-grained feature map.

[0035] S2133: Design two groups of dynamic weighting functions to fuse the fine-grained features and the coarse-grained features respectively; one group of branches fuses the fine-grained features to obtain fine-grained weighted features F fw , so that the final fusion features obtained have a larger proportion of fine-grained features and a smaller proportion of coarse-grained features; the other group of branches fuses the coarse-grained features to obtain coarse-grained weighted features F cw , so that the final fusion features obtained have a larger proportion of coarse-grained features and a smaller proportion of fine-grained features.

[0036] Preferably, the calculation formula of the attention weight W i of the i-th channel of the channel attention module in step S2131 is:

[0037]

[0038] Wherein, Conv1 and Conv2 represent convolution layers with a convolution kernel size of 1x1, δ represents a LeakyReLU activation function, σ represents a Sigmoid activation function, represents global average pooling, represents global maximum pooling.

[0039] Preferably, the formula of the fine-grained feature dynamic weighting function F fw in step S2133 is:

[0040] F fw =α⊙(F f ⊙W f )+(1-α)⊙(F c ⊙W c );

[0041] Fcw = (1 - a) O (F f O W f ) + a O (F c O W c ) ;

[0042]

[0043] coarse-grained feature dynamic weighting function F cw The formula is:

[0044] F fw = b O (F f O W f ) + (1 - b) O (F c O W c ) ;

[0045] F cw = (1 - b) O (F f O W f ) + b O (F c O W c ) ;

[0046]

[0047] wherein H(·) represents an entropy for measuring the information size of the feature map; O represents element-wise multiplication;

[0048] represents the proportion of fine-grained information in the total information;

[0049] represents the proportion of coarse-grained information in the total information;

[0050] a and b are two scalars for adjusting the proportion of feature fusion;

[0051] Because F f , W f , F c , and W c are updated with the training of the model, F fw and F cw are regarded as dynamic weights, and are set as the lower bounds of a and b, respectively. Preferably, the specific steps of the step S214 are:

[0052] Preferably, the specific steps of the step S214 are:

[0053] S2141: Cascade the fine-grained weighted feature F fw and the coarse-grained weighted feature F cw obtained in the step S2133, and add them element-wise to the part not selected in the step S211, and then cascade the selected part in the step S211 to obtain a merged feature map.

[0054] S2142: Each module in the reversible network contains the same double-branch dynamic weighting component (from double-branch feature extraction to dynamic weighting function), and the merged feature map is sent to the next module to perform the same operation until all modules are traversed to obtain the final feature maps of the content image and the style image.

[0055] Preferably, the merged feature map in the step S2142 is converted into a stylized feature map by a linear transformation module and is used to generate an image of a target color in the reverse inference process.

[0056] Preferably, the loss function in the step S22 includes three terms, specifically:

[0057] Matting Laplacian loss function L m , the formula is:

[0058]

[0059] Where N represents the number of pixels of the image, V c [I cs ] is the result of vectorizing the stylized image I cs in the cth channel, and M is the Matting Laplacian matrix of the content image I c ; T is the transpose matrix;

[0060] Style loss function L s , the formula is:

[0061]

[0062] Where I s represents the style image, I cs represents the stylized image; l is the number of layers; represents the i-th layer of the VGG-19 network, and μ and σ represent the mean and variance of the feature map, respectively;

[0063] Cycle consistency loss L cyc is calculated using the L1 norm, and the formula is:

[0064]

[0065] Where I c is the content image, is the reconstructed content image obtained by inverse recovery of the stylized image.

[0066] Compared with the prior art, the present application has the following characteristics and beneficial effects:

[0067] (1) The double-branch feature extraction structure and visual simulation operation are introduced, in the forward reasoning process, the fine-grained and coarse-grained features in the image can be effectively captured, and the dynamic weighting function is designed to fuse the fine-grained and coarse-grained features respectively; through the double-branch dynamic weighting component, the glass color transfer is more effectively realized;

[0068] (2) For the industrial glass image with small data amount and high data similarity, the effective color transfer can also be realized, and the problem of difficult sample collection of glass defect images under different light sources is solved. BRIEF DESCRIPTION OF DRAWINGS

[0069] Figure 1 It is a flowchart of the glass color transfer method based on dynamic weight of the application;

[0070] Figure 2 It is a model diagram of the glass color transfer method based on dynamic weight of the application;

[0071] Figure 3 It is a dynamic weight module diagram of the glass color transfer method based on dynamic weight of the application;

[0072] Figure 4 It is a schematic diagram of a black mobile phone screen glass defect image sample of the glass color transfer method based on dynamic weight of the application;

[0073] Figure 5 It is a comparison diagram of different glass color transfer models in the glass color transfer method based on dynamic weight of the application; wherein, (a) is the effect diagram generated by the glass color transfer model proposed in the application; (b) is the glass color transfer effect diagram based on the AesPA-Net model; (c) is the glass color transfer effect diagram based on the MicroAST model; (d) is the glass color transfer effect diagram based on the ArtFlow model; through comparison, the differences of color transfer of different models can be directly seen, and the application realizes the optimal result in color transfer and content consistency. DETAILED DESCRIPTION

[0074] To further describe the technical scheme disclosed by the application, the following further description is made in conjunction with the drawings of the specification. Embodiment: as shown in the figure, the glass color transfer method based on dynamic weight comprises the following steps: Figure 1

[0075] S1 Constructing a data set: collect data, construct a data set, and pretreat the data in the data set, adjust the image to a uniform size, and divide it into a training set and a test set;

[0076] ​S2, constructing and training the model: based on the reversible network, introducing a double-branch feature extraction structure and visual simulation operation to construct the model; and training the model to obtain the trained dynamic weight-based glass color transfer model;

[0077] The specific steps of the step S2 are:

[0078] S21, constructing the model, as shown in Figures 2-3 The specific steps are:

[0079] S211, introducing a double-branch feature extraction structure composed of two convolution kernels with different sizes to extract fine-grained features and coarse-grained features at the same time;

[0080] The specific steps of the step S211 are:

[0081] S211, in the forward inference process, input the content image and the style image at the same time, and after the padding unit, divide them into two parts along the channel dimension, one of which is input into the projection network composed of 3 convolution layers to obtain the feature map;

[0082] S212, using a convolution kernel with a size of 3x3 and a convolution kernel with a size of 5x5 to process the feature map obtained in step S211 to obtain the corresponding fine-grained feature map and coarse-grained feature map respectively; that is, forming two branches, one branch is to extract the fine-grained feature map using a convolution kernel with a size of 3x3, and the other branch is to extract the coarse-grained feature map using a convolution kernel with a size of 5x5;

[0083] S212, then through the upsampling unit and the downsampling unit for visual simulation;

[0084] In the step S212, the fine-grained feature map is passed through the upsampling unit to simulate the visual process of "looking closely"; the coarse-grained feature map is passed through the downsampling unit to simulate the visual process of "looking from a distance"; and the key features are emphasized; the specific steps are:

[0085] S2121, using interpolation and a convolution kernel with a size of 3x3 to upsample the extracted fine-grained feature map F f and enlarge it by two times to simulate the visual process of "looking closely", and the formula is:

[0086] F fu =Conv(Interpolate(F f ,scale_factor=2));

[0087] Where, F fuF is a fine-grained simulated visual feature map; Conv represents convolution; Interpolate represents difference up-sampling; scale_factor=2 represents 2 times magnification;

[0088] S2122: The extracted coarse-grained feature map F c is down-sampled and reduced by two times to simulate the visual process of human "looking from a distance", and the formula is:

[0089] F cd =Conv(F c , stride=2);

[0090] Wherein, F cd is a coarse-grained simulated visual feature map; Conv represents convolution; stride=2 represents 2 times reduction;

[0091] S213: Introducing a channel attention unit to give two branches weights respectively, and designing a dynamic weight function to fuse the extracted fine-grained features and coarse-grained features respectively;

[0092] As shown in Figure 3 , the specific steps of the step S213 are:

[0093] S2131: Introducing a channel attention module, which is composed of a pooling layer, a convolution layer and an activation function;

[0094] S2132: The fine-grained feature map F f , the fine-grained simulated visual feature map F fu , the coarse-grained feature map F c and the coarse-grained simulated visual feature map F cd are sent into the channel attention module respectively, and the channel attention weight is calculated to obtain four corresponding channel attention weights;

[0095] The channel attention weights of the fine-grained feature map F f and the fine-grained simulated visual feature map F fu are added to obtain the final channel attention weight W f of the fine-grained feature map;

[0096] The channel attention weights of the coarse-grained feature map F c and the coarse-grained simulated visual feature map F cd are added to obtain the final channel attention weight W c of the coarse-grained feature map;

[0097] S2133: design two groups of dynamic weighting functions to fuse the fine-grained features and the coarse-grained features respectively; one group of functions fuses the fine-grained branches to obtain fine-grained weighted features F fw , so that a larger proportion of fine-grained features and a smaller proportion of coarse-grained features are obtained in the final fusion features; the other group of functions fuses the coarse-grained branches to obtain coarse-grained weighted features F cw , so that a larger proportion of coarse-grained features and a smaller proportion of fine-grained features are obtained in the final fusion features;

[0098] The calculation formula of the attention weight W i of the i-th channel of the channel attention module in step S2131 is:

[0099]

[0100] Wherein, Conv1 and Conv2 represent convolution layers with a convolution kernel size of 1x1, δ represents a LeakyReLU activation function, σ represents a Sigmoid activation function, represents global average pooling, represents global maximum pooling;

[0101] The formula of the fine-grained feature dynamic weighting function F fw in step S2133 is:

[0102] F fw = α ⊙ (F f ⊙ W f ) + (1-α) ⊙ (F c ⊙ W c ) ;

[0103] F cw = (1-α) ⊙ (F f ⊙ W f ) + α ⊙ (F c ⊙ W c ) ;

[0104]

[0105] The formula of the coarse-grained feature dynamic weighting function F cw is:

[0106] F fw = β ⊙ (F f ⊙ W f ) + (1-β) ⊙ (F c ⊙ W c ) ;

[0107] F cw = (1-β) ⊙ (F f ⊙ Wf )+β⊙(F c ⊙W c );

[0108]

[0109] where H(·) denotes entropy, which is used to measure the information size of the feature map; ⊙ represents element-wise multiplication; s.t. is the abbreviation of subject to, and the content after it is the constraint condition;

[0110] represents the proportion of fine-grained information in the total information;

[0111] represents the proportion of coarse-grained information in the total information;

[0112] α and β are two scalars, which are used to adjust the proportion of feature fusion;

[0113] Because F f , W f , F c , and W c are updated with the training of the model, F and W are regarded as dynamic weights, and are set as the lower bound of α and β, respectively;

[0114] S214: Finally, the fine-grained and coarse-grained features processed by the dynamic weight function are concatenated;

[0115] The specific steps of the step S214 are as follows:

[0116] S2141: The fine-grained weighted feature F fw and the coarse-grained weighted feature F cw obtained in the step S2133 are concatenated, and are added element-wise with the part not selected in the step S211, and then are concatenated with the part selected in the step S211 to obtain the merged feature map;

[0117] S2142: Each module in the reversible network contains the same double-branch dynamic weighting component (from the double-branch feature extraction to the dynamic weighting function), and the merged feature map is sent to the next module to perform the same operation until all the modules are traversed to obtain the final feature maps of the content image and the style image;

[0118] The merged feature map in the step S2142 is converted into a stylized feature map through a linear transformation module, and is used to generate an image of a target color in the reverse inference process;

[0119] S22 training model: input the training set for training, iterate the model through the loss function, obtain the trained glass color transfer model based on dynamic weight, and then test the model performance on the test set; the loss function in the step S22 includes three items, specifically:

[0120] Matting Laplacian loss function L m , the formula is:

[0121]

[0122] Wherein, N represents the number of pixels of the image, V c [I cs ] is the result of the vectorization of the stylized image I cs The cth channel, M is the Matting Laplacian matrix of the content image I c ; T is the transpose matrix;

[0123] Style loss function L s , the formula is:

[0124]

[0125] Wherein, I s Indicates the style image, I cs Indicates the stylized image; l is the number of layers; Indicates the i-th layer of VGG-19 network, and μ and σ respectively represent the mean and variance of the feature map;

[0126] Cycle consistency loss L cyc , calculated using L1 norm, the formula is:

[0127]

[0128] Wherein, I c For content image, For the reconstructed content image obtained by inverse recovery of the stylized image;

[0129] S3 output result: input the style image and the content image into the glass color transfer model based on dynamic weight trained in step S2, to generate the image of target color. As Figure 4 The schematic diagram of the black mobile phone screen glass defect image sample based on the glass color transfer method based on dynamic weight is shown.

[0130] The glass color transfer model based on dynamic weight of the application is compared with other glass color transfer models (AesPA-Net model, MicroAST model and ArtFlow model), as Figure 5Fig. 1 shows a comparison of different glass color transfer models in terms of color transfer effect in the dynamic weight-based glass color transfer method of the present application; wherein, Figure 5 Fig. 1(a) shows the effect diagram generated by the dynamic weight-based glass color transfer model of the present application; Figure 5 Fig. 1(b) shows the glass color transfer effect diagram based on the AesPA-Net model; Figure 5 Fig. 1(c) shows the glass color transfer effect diagram based on the MicroAST model; Figure 5 Fig. 1(d) shows the glass color transfer effect diagram based on the ArtFlow model.

[0131] AesPA-Net proposes a pattern repeatability measurement method to quantify the repetition degree of patterns in style images, and accordingly finds the best balance point between local and global style expression to improve the quality of stylized images. However, it is difficult to accurately find this best point in practical applications, which may result in loss of details. MicroAST designs two micro-encoders to extract content and style features, respectively, and a micro-decoder to generate the final image. Limited by the limited expression ability of the micro-network, the synthesized image is prone to lose details. ArtFlow uses reversible neural flow to realize style transfer, which maps the content image and the style image to the latent space through forward inference, and then uses WCT or AdaIN for style transfer in the reverse inference process. Since the reversible flow structure may introduce redundant information, the generated image may lose details.

[0132] From the comparison of Figure 5 , it can be seen intuitively that the different models have different color transfer differences. The present application introduces a double-branch feature extraction structure and a visual simulation operation, which can effectively capture fine-grained and coarse-grained features in the image during the forward inference process, and designs a dynamic weighting function to fuse fine-grained and coarse-grained features, respectively. Through upsampling and downsampling, the visual process of "looking closely" and "looking far" is simulated, and key features are emphasized. Optimal results are achieved in terms of color transfer and content consistency.

[0133] The quantitative evaluation results of the present application and other color transfer models are shown in Table 1. Structural similarity (SSIM) is used to evaluate content consistency, and Gram loss is used to evaluate color transfer effect. The closer the SSIM value is to 1, the more similar the two images are in content. The lower the Gram loss value, the closer the color of the generated image is to the target image. Under the same training conditions, the present application achieves the highest SSIM score and the lowest Gram loss score, indicating that the present application can accurately transfer colors while maintaining the consistency of the content to the greatest extent.

[0134] Table 1 quantitative evaluation results of the present application and other color migration models

[0135] Method The present invention AesPA-Net MicroAST ArtFlow SSIM↑ 0.685 0.507 0.574 0.264 Gram loss↓ 2.335 3.180 3.799 2.468

[0136] From the evaluation results in Table 1, compared with other technologies, the present application introduces a double-branch feature extraction structure and a visual simulation operation, which can effectively capture fine-grained and coarse-grained features in the image during the forward reasoning process, and a dynamic weighting function is designed to fuse fine-grained and coarse-grained features respectively; can generate a glass defect image of the corresponding color according to the target image, and keep the content consistency of the generated image and the content image, can more effectively realize the glass color migration, solve the problem of difficult sample collection of glass defect image under different light sources, and has higher application value.

[0137] The above examples are only specific implementations of the present application, and are not a limitation of the present application. Any modification, replacement or improvement made by those skilled in the art without departing from the core idea of the present application through conventional technical means shall fall within the protection scope of the present application.

Claims

1. A dynamic weight based glass color migration method, characterized in that, The method comprises the following steps: S1, constructing a data set: collecting data, constructing a data set, and preprocessing data in the data set; S2, constructing and training a model: improving based on a reversible network, introducing a double-branch feature extraction structure and a visual simulation operation to construct a model, and training the model to obtain a trained dynamic weight-based glass color transfer model; S3, outputting a result: inputting a style image and a content image into the dynamic weight-based glass color transfer model trained in step S2 to generate an image of a target color. The specific steps of step S2 are:

2. The dynamic weight based glass color migration method of claim 1, wherein, S21, constructing a model, specifically: S211, introducing a double-branch feature extraction structure composed of two convolution kernels with different sizes to extract fine-grained features and coarse-grained features at the same time; S212, performing visual simulation through an up-sampling unit and a down-sampling unit, respectively; S213, introducing a channel attention unit to assign weights to the two branches, and designing a dynamic weight function to fuse the extracted fine-grained features and coarse-grained features, respectively; S214, finally performing a cascading operation on the two branches; S22, training a model: inputting a training set to train the model, iterating the model through a loss function, obtaining a trained dynamic weight-based glass color transfer model, and then testing the model performance on a test set. The specific steps of step S211 are:

3. The dynamic weight based glass color migration method of claim 2, wherein, S211, in the forward inference process, inputting a content image and a style image at the same time, and after passing through a padding unit, dividing them into two parts along the channel dimension, one of which is input into a projection network to obtain a feature map; S212, using a convolution kernel with a size of 3x3 and a convolution kernel with a size of 5x5 to process the feature map obtained in step S211, respectively obtaining a fine-grained feature map and a coarse-grained feature map. In step S212, the fine-grained feature map is simulated through an up-sampling unit to simulate the human "close look" visual process; the coarse-grained feature map is simulated through a down-sampling unit to simulate the human "distant look" visual process; the specific steps are:

4. The dynamic weight based glass color migration method of claim 2, wherein, The specific steps of step S213 are: S2121: The extracted fine-grained feature map F f Upsampling is performed and it is enlarged by two times to simulate the visual process of human "close look", and the formula is: F fu = Conv(Interpolate(F f , scale_factor = 2)); where F fu is a fine-grained simulation visual feature map; Conv represents convolution; Interpolate represents difference up-sampling; scale_factor = 2 represents 2 times magnification; S2122: The extracted coarse-grained feature map F c Down-sampling and reducing it by two times, simulating the visual process of human "looking from afar", the formula is: F cd = Conv(F c , stride = 2); where F cd is a coarse-grained simulation visual feature map; Conv represents convolution; stride=2 means reducing by 2 times.

5. The dynamic weight based glass color migration method of claim 3, wherein, S2131, introducing a channel attention module, which is composed of a pooling layer, a convolution layer and an activation function; Where H(·) represents entropy, which is used to measure the information size of the feature map; and represents element-wise multiplication; S2132: the fine-grained feature map F f , the fine-grained simulated visual feature map F fu , the coarse-grained feature map F c , and the coarse-grained simulated visual feature map F cd are respectively input into the channel attention module, and the channel attention weights are calculated to obtain four corresponding channel attention weights; The fine-grained feature map F f and the fine-grained simulated visual feature map F fu The calculated channel attention weights are added to obtain the final channel attention weight W of the fine-grained feature map f ; The coarse-grained feature map F is obtained c and the coarse-grained simulation visual feature map F cd The calculated channel attention weights are added to obtain the final channel attention weight W of the coarse-grained feature map c ; S2133: design two groups of dynamic weighting functions to perform the fusion of the fine-grained features and the coarse-grained features respectively; one group of functions processes the fine-grained branches to obtain fine-grained weighted features F fw , and the other group of functions processes the coarse-grained branches to obtain coarse-grained weighted features F cw .

6. The dynamic weight based glass color migration method of claim 5, wherein, The calculation formula of the attention weight W of the i-th channel of the channel attention module in the step S2131 is as follows: i The calculation formula of the attention weight W of the i-th channel of the channel attention module in the step S2131 is as follows: wherein Conv1 and Conv2 represent convolution layers with a kernel size of 1x1, δ represents a LeakyReLU activation function, and σ represents a Sigmoid activation function, represents a global average pooling, represents a global max pooling.

7. The dynamic weight based glass color migration method of claim 5, wherein, The fine-grained feature dynamic weighting function F in the step S2133 fw The formula is: F fw = a o (F f o W f ) + (1 - a) o (F c o W c ); F cw = (1 - a) O (F f O W f ) + a O (F c O W c ); Coarse-grained feature dynamic weighting function F cw The formula is: F fw = β o (F f o W f ) + (1 - β) o (F c o W c ); F cw = (1 - β) O (F f O W f ) + β O (F c O W c ); α and β are two scalars used to adjust the proportion of feature fusion; represents the proportion of fine-grained information to the total information; represents the proportion of coarse-grained information to the total information; The specific steps of step S214 are: Because F f , W f , F c , and W c are updated as the model is trained, they are treated as dynamic weights and are set to the lower bounds of a and b, respectively. and are treated as dynamic weights and are set to the lower bounds of a and b, respectively.

8. The dynamic weight based glass color migration method of claim 5, wherein, S2142, each module in the reversible network contains the same double-branch dynamic weighting component, and the merged feature map is sent to the next module to perform the same operation until all modules are traversed to obtain the final feature maps of the content image and the style image. S2141: concatenate the fine-grained weighted feature F fw and the coarse-grained weighted feature F cw and the unselected part in step S211 element by element, and then concatenate the selected part in step S211, i.e., the part input into the projection network, to obtain a merged feature map; In step S2142, the merged feature map is converted into a stylized feature map through a linear transformation module, and is used to generate an image of a target color in the backward inference process.

9. The dynamic weight based glass color migration method of claim 8, wherein, The loss function in step S22 includes three terms, specifically:

10. The dynamic weight based glass color migration method of claim 8, wherein, ​ Matting Laplacian loss function L m , the formula is: where N represents the number of pixels of the image, V c [I cs ] is the stylized image I cs resulted from the vectorization of the cth channel, M is the Matting Laplacian matrix of the content image I c ; T is the transpose matrix; Style loss function L s , the formula is: where I s denotes the stylized image, I cs denotes the stylized image; l is the number of layers; denotes the i-th layer of the VGG-19 network, μ and σ denote the mean and variance of the feature map, respectively; cycle consistency loss L cyc using the L1 norm, with the formula: where I c is the content image, is the stylized image inverse-restore to get the reconstructed content image.

Citation Information

Patent Citations

  • Glass image color migration system, migration method, device and medium

    CN119810229A