Industrial product accumulated damage image generation method based on improved DCGAN
The IPD-GAN model stabilizes training and enhances image quality by using a gradient penalty in the Wasserstein distance loss and residual blocks, effectively generating high-resolution industrial product damage images.
Patent Information
- Application Number
- CN202510378388.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-15
AI Technical Summary
Existing GAN and its variant networks are prone to problems such as gradient disappearance, training instability and poor image quality when generating high-resolution images, and traditional data enhancement methods cannot effectively expand industrial product damage data sets.
Based on the DCGAN model, the network structure is optimized to generate high-resolution industrial product damage images by introducing Wasserstein distance loss function, L1 loss and SSIM loss with gradient penalty terms, and adding residual blocks to the generator.
The expansion of industrial damage sample data sets is achieved, the generation image quality and diversity are improved, the feature distribution is closer to the real image, solving the problems of training instability and gradient disappearance, and the generator performance is improved.
Smart Images

Figure CN120318352A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image generation technology, and in particular to a method for generating cumulative damage images of industrial products based on an improved DCGAN. Background Art
[0002] In order to study the impact of the atmospheric environment on industrial products, my country has established more than 10 atmospheric natural environment test stations, corresponding to various types of climates such as high temperature, high humidity, high salt, high cold, strong solar radiation, acid rain, low humidity, and low temperature. In view of the common high temperature, rainy fog, acid precipitation and other phenomena in Chongqing, which cause industrial product parts to be corroded, and then affect product precision and production efficiency, and may even cause accidents such as explosions and combustion, Chongqing University, together with the 59th Institute of China Ordnance Industry and Chongqing Technology and Business University, jointly carried out research to form a set of autonomous inspection and intelligent evaluation robot systems and related technical systems for atmospheric natural environment tests.
[0003] In the process of research, a large number of industrial product atmospheric natural environment test data sets are often required to complete product defect recognition and classification tasks using deep learning models. The training of the models depends on the use of a large number of labeled data sets. In the research on the environmental adaptability of industrial products, there are problems such as small data volume and limited samples, and data augmentation is required to obtain sufficient data sets. Traditional data augmentation methods include rotating, mirroring, translating samples, and adding noise to generate new samples. Since the augmented samples have a high similarity with the original samples, it is impossible to ensure that the newly generated samples are beneficial to model training, and sometimes it even exacerbates the overfitting degree of the model. At present, the method of data set augmentation adopted by many researchers is the generative adversarial network GAN method. As a generative model, the generative adversarial network can mine useful feature information from real images, thereby generating fake data similar to real samples. Goodfellow et al. proposed the most primitive generative adversarial network model GAN, which introduced random noise to imitate real samples and finally generated new samples, solving the problem of unbalanced original samples. However, this model has unstable training, difficult convergence, and problems such as gradient explosion and low image generation quality. Many researchers have improved and optimized based on the GAN model, and then many new models have emerged. MIRZA et al. proposed the conditional generative adversarial network model CGAN, adding conditions to the network to make the network generate samples in a given direction. Radford et al. proposed the deep convolutional generative adversarial network model DCGAN, introducing the idea of convolutional neural network CNN into the generative adversarial network, solving the problem of blurred generated images in the original network. Shi Hongyu et al. proposed an improved ACGAN strip steel small sample data augmentation method, designing a denoising structure in the network and introducing the idea of cascade fusion, improving the clarity and accuracy of the generated images, but still need to improve the training speed. Han Xiang et al. proposed a defect sweet cherry image enhancement method based on DCGAN, introducing a multi-scale residual block (MSRB) and CBAM attention mechanism in the generator, enhancing the model's feature expression ability and the detail quality of the generated images, and improving the gradient flow at the same time.
[0004] Disadvantages of the prior art: Most of the existing GANs and their variant networks are used to generate images with a resolution of 128×128 and below. When the resolution of the generated image reaches 512×512 and the number of network layers is deepened to a certain extent, problems such as gradient disappearance, unstable training, and poor quality of the generated images will occur. Therefore, it cannot effectively generate high-resolution images. Using progressive DCGAN training, starting from training a shallow network that generates low-resolution images, the network can gradually adapt to more complex tasks, but it greatly lengthens the training time and the implementation process is too complex. Summary of the Invention
[0005] An industrial product cumulative damage image generation method based on an improved DCGAN provided by the present invention optimizes problems such as unstable network model training and easy occurrence of gradient disappearance when the DCGAN generates high-resolution images.
[0006] To achieve the above object, a key aspect of an industrial product cumulative damage image generation method based on an improved DCGAN provided by the present invention includes the following steps:
[0007] Step 1: Based on the DCGAN model framework, construct an industrial product cumulative damage image generation network IPD-GAN based on the improved DCGAN. The industrial product cumulative damage image generation network IPD-GAN is provided with a generator and a discriminator;
[0008] Step 2: The generator obtains random noise, processes the random noise, generates a damage image, and transmits it to the discriminator;
[0009] Step 3: The discriminator obtains a real image, discriminates the authenticity of the real image and the damage image, uses the Wasserstein distance loss function with a gradient penalty term as the model adversarial loss according to the discrimination result, and combines the L1 loss and the SSIM loss to guide the training of the generator, comprehensively optimizing the system performance;
[0010] Step 4: Use the trained industrial product cumulative damage image generation network IPD-GAN based on the improved DCGAN to generate damage images.
[0011] Through the above design, an improvement is made on the basis of the DCGAN model. The loss function is modified to use the Wasserstein distance with a gradient penalty term as the adversarial loss, and the L1 loss and the SSIM loss are added to guide the training of the generator, optimizing problems such as unstable training and gradient disappearance existing in the DCGAN, and improving the network's learning of product damage defect details and structural integrity; realizing the expansion of the industrial damage sample dataset.
[0012] The industrial product cumulative damage image generation network IPD-GAN based on the improved DCGAN has a significant improvement in the overall distribution and richness of the generated samples. The feature distribution is closer to the real image, and the quality of the generated images is higher, reflecting the reliability of the present invention for generating industrial product damage data under the atmospheric environment.
[0013] Preferably: In the step 1, the generator is provided with a first network layer, a second network layer, a third network layer, a fourth network layer, a fifth network layer, a sixth network layer, a seventh network layer, and an eighth network layer connected in sequence;
[0014] The first network layer, the second network layer, the third network layer, the fourth network layer, the fifth network layer, the sixth network layer, and the seventh network layer are all provided with a transposed convolutional layer and a batch normalization layer connected in sequence. A ReLU activation function is provided at the backend of the batch normalization layer; a residual block is further connected after the activation function of the second network layer, the third network layer, and the sixth network layer; the eighth network layer is provided with a transposed convolutional layer and a Tanh activation function.
[0015] Since the generated image has a relatively high resolution, in order to enable the generation network to generate higher-quality images, the generator is provided with a total of eight layers of structures, which are used to generate a three-channel image of 512×512 from a random noise with an input dimension of 100. First, a transposed convolution of the first structure is performed to generate a feature map with a higher dimension. The size of the convolutional kernel is 4, the stride is 1, and the number of channels of the feature map is 4096. Subsequently, a batch normalization layer and a ReLU activation function are connected behind the convolutional layer to accelerate training, stabilize gradient propagation, and increase the network non-linearity. The operations of the first structure layer are repeated in the subsequent several layers. The spatial dimension of the feature map is gradually increased and the number of channels of the feature map is decreased through transposed convolution, proceeding from 4096×4×4 to 2048×8×8 to 1024×16×16 until finally an image of 3×512×512 is output.
[0016] Compared with the traditional DCGAN network, the network of the present invention has a deeper network hierarchy. In order to overcome the problem that deep networks are prone to vanishing gradients, reduce the network computational complexity, and at the same time ensure the stable training of the network, residual blocks are added after the activation functions of the second layer, the third layer, and the sixth layer of the generator. The introduction of the residual block can not only increase the stability of network training but also help the generator effectively retain low-level feature information, enabling it to learn more abundant information and generate higher-quality images.
[0017] Preferably: The residual block Res is provided with a convolutional layer, a batch normalization layer, a ReLU activation function, a convolutional layer, a batch normalization layer, a residual connection connected in sequence, and outputs after passing through a ReLU activation.
[0018] Preferably: In the step 1, the discriminator is provided with a first network block, a second network block, a third network block, a fourth network block, a fifth network block, a sixth network block, a seventh network block, and an eighth network block connected in sequence;
[0019] The first network block is provided with a convolutional layer, and a LeakyReLU activation function is provided at the backend of the convolutional layer;
[0020] The second network block, the third network block, the fourth network block, the fifth network block, the sixth network block, and the seventh network block are provided with a convolutional layer and a batch normalization layer connected in sequence, and a LeakyReLU activation function is provided at the backend of the batch normalization layer; a dropout layer is connected after the LeakyReLU activation functions of the sixth network block and the seventh network block;
[0021] The eighth network block is provided with a convolutional layer.
[0022] The discriminator also consists of an 8-layer structure. The discriminator adopts 8 convolutional layers. First, it receives an image of 512×512×3 as an input, and uses a convolutional kernel of 4×4 for convolution with a stride of 2, and the number of output feature map channels is 32. The number of channels of the feature map is gradually increased through multiple convolutions to extract the high-level features of the image. A batch normalization layer is connected after the convolutional layer in the middle structure layer to standardize the output of the convolution, help accelerate training and stabilize the network, and prevent the problems of gradient disappearance and gradient explosion; then the LeakyReLU activation function is used to help the gradient propagation during the training process, and the negative slope of the LeakyReLU activation function is set to 0.2.
[0023] Since the amount of data in the industrial product damage dataset used for training is small, in order to reduce the overfitting of the network and increase the diversity of the generated samples, a dropout layer is added after the sixth and seventh structural levels, which can not only prevent the premature discarding of feature information from resulting in poor learning effects, but also reduce overfitting, and the random inactivation rate is set to 0.3.
[0024] The traditional DCGAN discriminator network uses binary cross-entropy loss. In order to make the discriminator loss a probability value between 0 and 1, a Sigmoid activation function is used for output in the last layer. However, in the present invention, the Wasserstein distance is used to calculate the difference between the generated samples and the real samples, so this activation function is removed and the real value is directly output.
[0025] Preferably: in the step 3, the Wasserstein distance is used to measure the "distance" between the real samples and the generated samples, so that the training process is more stable. The goal of the generator G is to minimize the distance between the generated samples and the real samples, while the discriminator, on the contrary, needs to maximize this distance; the expression of the Wasserstein distance loss function is:
[0026]
[0027] Among them, W(P r ,P g ) is the Wasserstein distance, P r and P g are two probability distributions, Π(Pr ,P g ) represents P and P g The set of all possible joint distribution probabilities; inf is the lower bound; x is the real sample; y is the generated sample; ||xy|| is the distance between x and y; E represents the expectation, and γ is the joint distribution;
[0028] Since the Wasserstein distance requires the discriminator to satisfy the Lipschitz continuity condition, the gradient norm must be less than or equal to 1, otherwise it is easy to cause the gradient to vanish or explode. Therefore, the gradient penalty term GP is introduced in the present invention to satisfy the Lipschitz continuity constraint condition, making the training process more stable. The gradient penalty term L is introduced into the Wasserstein distance loss function. GP The expression is:
[0029]
[0030] Among them, λ is the penalty coefficient; Represents the linear interpolation sample between the generated sample and the real sample; is the gradient of the discriminator with respect to the input sample; ||·||2 is the L2 norm.
[0031] The original DCGAN uses cross entropy loss, which calculates the binary classification error between the generated image and the real image. This loss mainly promotes the image generated by the generator to be realistic enough to deceive the generator, but cannot clearly guide the generator on how to generate realistic and diverse pictures. There are problems such as mode collapse, unstable training, and insufficient diversity of generated images. The present invention uses Wasserstein distance with gradient penalty as adversarial loss, and adds L1 loss and SSIM loss to comprehensively optimize the authenticity and detail features of the generated image.
[0032] As a preference: in step 3, the L1 loss L L1 The expression is as follows:
[0033]
[0034] Among them, ||·||1 is the L1 norm;
[0035] The SSIM loss L SSIM The expression is as follows:
[0036] L SSIM =1-SSIM(y,x)
[0037] SSIM(y,x)=l(x,y)·c(x,y)·s(x,y)
[0038] Among them, l(x, y) is the brightness of the image, c(x, y) is the contrast of the image, and s(x, y) is the structure of the image; the value of SSIM ranges from 0 to 1, where 1 represents that the two images are exactly the same, and 0 represents that the two images are completely different. The closer the SSIM value is to 1, the better the quality of the generated image.
[0039]
[0040] Among them, μ represents the mean of the sample, σ represents the variance of the sample, and σ xy represents the covariance of the sample; C1 = (K1L) 2 and C2 = (K2L) 2 and C3 = C2 / 2; K1 = 0.01, K2 = 0.03, and L is the pixel value range. C1, C2, and C3 are all constants used to avoid the denominator being equal to 0.
[0041] The present invention requires the generated image to have a relatively high resolution. To improve the quality of image generation, it is necessary to optimize the output of the generator from multiple dimensions, enabling the network to not only observe the authenticity of the image but also learn and improve the details of product damage defects and structural integrity. Therefore, in addition to selecting the Wasserstein distance with a gradient penalty term as the adversarial loss, an L1 loss and an SSIM loss are also added to guide the training of the generator.
[0042] The L1 loss measures the pixel-level difference between the real image and the generated image, and the goal is to minimize this difference to keep the generated image with details and clarity. The SSIM loss considers the structural similarity of the image by calculating the similarity in three aspects: the brightness l(x, y), contrast c(x, y), and structure s(x, y) of the image.
[0043] Preferably, the objective formula L D of the loss function of the discriminator is expressed as follows:
[0044]
[0045] The objective formula L G of the loss function of the generator is expressed as follows:
[0046]
[0047] Among them, represents the expectation of sample x in the discriminator, represents the expectation of sample y in the discriminator; λ L1 is the weight coefficient of the L1 loss, and λ SSIM is the weight coefficient of the SSIM loss, used to balance the influence of the loss terms on the generator.
[0048] Advantages of the present invention: The dataset expansion of industrial damage samples is achieved. Based on the DCGAN model, improvements are made. The loss function is modified to the Wasserstein distance with a gradient penalty term as the adversarial loss. The L1 loss and SSIM loss are added to guide the training of the generator, optimizing problems such as unstable training and mode collapse existing in DCGAN. Residual blocks are added to the generator to alleviate the problem of vanishing gradients in deep networks for generating high-resolution images, while accelerating the training convergence and improving the performance of the generator. There is a significant improvement in the overall distribution and richness of the samples generated by IPD-GAN. The feature distribution is closer to real images, and the quality of the generated images is higher, reflecting the reliability of the present invention for generating industrial product damage data under atmospheric conditions. Description of the Drawings
[0049] Figure 1 It is the IPD-GAN network structure diagram in the embodiment;
[0050] Figure 2 It is the generator structure diagram of IPD-GAN in the embodiment;
[0051] Figure 3 It is the residual block structure diagram in the embodiment;
[0052] Figure 4 It is the discriminator structure diagram of IPD-GAN in the embodiment;
[0053] Figure 5 It is the damaged image generated by WGAN-GP in the embodiment;
[0054] Figure 6 It is the damaged image and the original damaged image generated by progressive DCGAN in the embodiment;
[0055] Figure 7 It is the damaged image and the original damaged image generated by IPD-GAN in the embodiment. Detailed Embodiment
[0056] The present invention will be further described in detail below with reference to the drawings and specific examples. The following examples or drawings are used to illustrate the present invention, but do not limit the scope of the present invention.
[0057] A method for generating industrial product cumulative damage images based on an improved DCGAN includes the following steps:
[0058] Step 1: Based on the DCGAN model framework, construct an industrial product cumulative damage image generation network IPD-GAN based on the improved DCGAN. The industrial product cumulative damage image generation network IPD-GAN is provided with a generator and a discriminator, as Figure 1 shown;
[0059] Step 2: The generator obtains random noise, processes the random noise to generate a damaged image, and transmits it to the discriminator;
[0060] Step 3: The discriminator obtains real images, discriminates the authenticity between the real images and the damaged images, uses the Wasserstein distance loss function with a gradient penalty term as the model adversarial loss according to the discrimination results, and combines the L1 loss and the SSIM loss to guide the training of the generator, comprehensively optimizing the system performance;
[0061] Step 4: Use the trained industrial product cumulative damage image generation network IPD-GAN based on the improved DCGAN to generate damaged images.
[0062] As Figure 2 shown, the generator is provided with a first network layer, a second network layer, a third network layer, a fourth network layer, a fifth network layer, a sixth network layer, a seventh network layer, and an eighth network layer connected in sequence;
[0063] The first network layer, the second network layer, the third network layer, the fourth network layer, the fifth network layer, the sixth network layer, and the seventh network layer are all provided with a transposed convolutional layer and a batch normalization layer connected in sequence, and a ReLU activation function is arranged at the backend of the batch normalization layer; a residual block is also connected after the ReLU activation function of the second network layer, the third network layer, and the sixth network layer; the eighth network layer is provided with a transposed convolutional layer, and a Tanh activation function is arranged at the backend of the transposed convolutional layer.
[0064] Since the generated image has a high resolution, in order to enable the generation network to generate images of high quality, the generator is provided with a total of eight layers of structures to generate a three-channel image of 512×512 from random noise with an input dimension of 100. First, a transposed convolution of the first structure is used to generate a feature map with a higher dimension. The size of the convolutional kernel is 4, the stride is 1, and the number of channels of the feature map is 4096. Subsequently, a batch normalization layer and a ReLU activation function are connected behind the convolutional layer to accelerate training, stabilize the gradient propagation, and increase the network non-linearity. The operations of the first structure layer are repeated in the subsequent several layers, and the spatial dimension of the feature map is gradually increased and the number of channels of the feature map is reduced through transposed convolution, proceeding from 4096×4×4 to 2048×8×8 to 1024×16×16 until finally an image of 3×512×512 is output.
[0065] Compared with the traditional DCGAN network, the network of the present invention has a deeper network hierarchy. In order to overcome the problem that deep networks are prone to gradient disappearance, reduce the network computing complexity, and at the same time ensure the stable training of the network, residual blocks are added after the activation functions of the second, third, and sixth layers of the generator. The introduction of the residual blocks can not only increase the stability of network training, but also help the generator effectively retain low-level feature information, enabling it to learn richer information and generate higher-quality images.
[0066] As Figure 3 shown, the residual block Res is provided with a convolutional layer, a batch normalization layer, a ReLU activation function, a convolutional layer, a batch normalization layer, a residual connection connected in sequence, and outputs through a ReLU activation.
[0067] As Figure 4 shown, the discriminator is provided with a first network block, a second network block, a third network block, a fourth network block, a fifth network block, a sixth network block, a seventh network block, and an eighth network block connected in sequence;
[0068] The first network block is provided with a convolutional layer, and a LeakyReLU activation function is arranged after the output end of the convolutional layer;
[0069] The second network block, the third network block, the fourth network block, the fifth network block, the sixth network block, and the seventh network block are provided with a convolutional layer and a batch normalization layer connected in sequence, and a LeakyReLU activation function is arranged at the backend of the batch normalization layer; a dropout layer is connected after the LeakyReLU activation functions of the sixth network block and the seventh network block;
[0070] The eighth network block is provided with a convolutional layer.
[0071] The discriminator also consists of an 8-layer structure. The discriminator adopts 8 convolutional layers. First, it receives an image of 512×512×3 as an input, and performs convolution with a 4×4 convolutional kernel, with a stride of 2, and the number of output feature map channels is 32. The number of channels of the feature map is gradually increased through multiple convolutions to extract high-level features of the image. A batch normalization layer is connected after the convolutional layer in the intermediate structure layer to standardize the output of the convolution, help accelerate training and stabilize the network, and prevent the problems of gradient disappearance and gradient explosion; then the LeakyReLU activation function is used to help the gradient propagation during the training process, and the negative slope of the LeakyReLU activation function is set to 0.2.
[0072] Since the data volume of the industrial product damage dataset used for training is small, in order to reduce the overfitting of the network and increase the diversity of the generated samples, a dropout layer is added after the sixth and seventh structural levels, which can not only prevent the premature discarding of feature information from resulting in poor learning effects, but also reduce overfitting, and the random inactivation rate is set to 0.3.
[0073] In the traditional DCGAN discriminator network, binary cross-entropy loss is used. In order to make the discriminator loss a probability value between 0 and 1, a Sigmoid activation function is used for output in the last layer. However, in the present invention, the Wasserstein distance is used to calculate the difference between the generated samples and the real samples, so this activation function is removed and the real-valued output is directly performed.
[0074] Figure 2-3 Among them, DeConv is the transposed convolution layer, ReLU is the ReLU activation function, BN is the batch normalization layer, Res is the residual block, Tanh is the Tanh activation function, Conv is the convolution layer, LeakyReLU is the LeakyReLU activation function, and Dropout is the dropout layer.
[0075] In the step 3, the expression of the Wasserstein distance loss function is:
[0076]
[0077] Among them, W(P r , P g ) is the Wasserstein distance, P r and P g are two probability distributions, Π(P r , P g ) represents the set of all possible joint distribution probabilities of P r and P g ; inf is to take the lower bound; x is the real sample; y is the generated sample; ||x - y|| is the distance between x and y; E represents the expectation, γ is the joint distribution;
[0078] A gradient penalty term is introduced into the Wasserstein distance loss function. The expression of the gradient penalty term L GP is:
[0079]
[0080] Among them, λ is the penalty coefficient; represents the linear interpolation sample between the generated sample and the real sample; is the gradient of the discriminator to the input sample; ||·||2 is the L2 norm.
[0081] In the step 3, the expression of the L1 loss L L1 is as follows:
[0082]
[0083] Among them, ||·||1 is the L1 norm;
[0084] The SSIM loss L SSIM has the following expression:
[0085] L SSIM = 1 - SSIM(y, x)
[0086] SSIM(y, x) = l(x, y)·c(x, y)·s(x, y)
[0087] where l(x, y) is the brightness of the image, c(x, y) is the contrast of the image, and s(x, y) is the structure of the image;
[0088]
[0089] where μ represents the mean of the sample, σ represents the variance of the sample, and σ xy represents the covariance of the sample; C1 = (K1L) 2 and C2 = (K2L) 2 and C3 = C2 / 2; K1 = 0.01, K2 = 0.03, and L is the pixel value range. C1, C2, and C3 are all constants used to avoid the denominator being equal to 0.
[0090] The objective formula L D of the loss function of the discriminator has the following expression:
[0091]
[0092] The objective formula L G of the loss function of the generator has the following expression:
[0093]
[0094] where represents the expectation of the sample x in the discriminator, represents the expectation of the sample y in the discriminator; λ L1 is the weight coefficient of the L1 loss, and λ SSIM is the weight coefficient of the SSIM loss, used to balance the influence of the loss terms on the generator.
[0095] Next, the performance of the present invention is verified through specific experiments.
[0096] The experimental running environment is shown in Table 1:
[0097] Table 1 Experimental running environment
[0098]
[0099] The dataset used in the experiment comes from industrial product damage images collected by the Chongqing Jiangjin Atmospheric Natural Environment Experimental Station. These images were cropped to a size of 512×512. Since the amount of data was too small, mirror operations were first performed on the images to expand the dataset size, and finally 134 images were obtained. Then, the network training parameters were set. The training batch size was set to 10, and each group of experiments was trained for 5000 epochs. Both the generator and the discriminator used the Adam optimization algorithm, with beta1 set to 0.7. To optimize the training process, a cosine annealing learning rate adjustment strategy was adopted. The initial learning rate of the discriminator was set to 0.0001, and the initial learning rate of the generator was set to 0.0002. The minimum learning rate of the discriminator was 0.00001, and the minimum learning rate of the generator was 0.00005. Taking the entire training length as one cycle, the gradient penalty term parameter lambda_GP = 10, and the weight coefficients of the L1 loss and the SSIM loss were set to lambda_L1 = 1 and lambda_SSIM = 0.1 respectively. To prevent the discriminator from quickly converging to 0, it was set that for each iteration, the generator was trained 3 times and the discriminator was trained 1 time.
[0100] To evaluate the images generated by the GAN network, multiple factors such as the fidelity, overall structural difference, pixel-level difference, and diversity of the images need to be considered. To objectively evaluate the effectiveness of the industrial product cumulative damage image generation network IPD-GAN based on the improved DCGAN and the image generation effect, this embodiment evaluates from the following three aspects.
[0101] (1) FID is an index for evaluating the distribution difference between the generated images and the real data. It measures the difference from the real data by calculating the Fréchet distance between the two distributions. Since FID takes into account the means and covariance matrices of the two distributions, it can better describe the distance between them, and the smaller the FID value, the closer the generated image is to the real image. The calculation formula of FID is as follows:
[0102]
[0103] where P and G respectively represent the sets of feature vectors of the real image distribution and the generated image distribution, μ P and μ G respectively represent the means of the sets of feature vectors of P and G, ∑p and ∑G are the covariance matrices of the sets of feature vectors of P and G respectively, and Tr represents the square root of the trace of the covariance matrix.
[0104] (2) PSNR measures the image reconstruction quality based on the mean square error MSE. The higher the value, the better the image quality and the closer it is to the actual distribution of the real features. The calculation formula of the mean square error (MSE) of RGB images is as follows:
[0105]
[0106] It represents the mean of the squared pixel-level differences between the generated image and the real image. The smaller the value, the smaller the difference between the generated image and the real image. I(i, j) represents the pixel of the real image at the pixel position (i, j), K(i, j) represents the pixel of the generated image at the pixel position (i, j), m and n represent the length and width of the image, and C represents the number of channels of the RGB image, where C = 3.
[0107] The calculation formula of PSNR is as follows:
[0108]
[0109] Among them, MAX represents the maximum possible value of the image pixels.
[0110] (3) IS is used to evaluate the quality and diversity of the generated images. IS classifies the generated images by using the pre-trained Inception v3 model and calculates the class distribution of the generated images. The calculation formula of IS is as follows:
[0111]
[0112] Among them, x ∼ p g represents generating pictures from the generator; p(y|x) represents inputting the generated picture x into Inception V3 to obtain a 1000-dimensional vector y, that is, the probability distribution of the picture belonging to each category; p(y) represents inputting the sample pictures selected from the generated images into Inception V3, respectively obtaining their own probability distribution vectors, and averaging these vectors to obtain the marginal distribution of all the pictures generated by the generator over all categories, as shown in the formula:
[0113]
[0114] DK represents calculating the KL divergence between p(y|x) and p(y). By calculating the KL divergence to measure the distance between two probability distributions, it is a non-negative value. The larger the value, the less similar the two probability distributions are. Generally speaking, as long as the distance between p(y|x) and p(y) is large enough, it can prove that this generation model is good enough. To more intuitively reflect the calculation of IS, in actual operation, the formula is rewritten as follows:
[0115]
[0116] First, calculate for the generated samples, and then calculate p(y|x (i) ) for each sample, and then calculate p(y|x (i) ) and the KL divergence, and finally calculate the exponent.
[0117] To evaluate the performance of different models, the dataset was trained on the IPD-GAN, DCGAN, and WGAN models respectively. The generation of industrial product damage samples at different iteration times under the progressive DCGAN, WGAN, and IPD-GAN training is shown in the following figures. Figure 5 Five images randomly selected from the images generated by the WGAN-GP model. It can be seen from the pictures that the generation effect of the WGAN-GP model is very unsatisfactory under the dataset of this embodiment, and the overall features and local pitting features of the images are very blurred. Figure 6 and Figure 7 Respectively show five generated images randomly selected at 3 iteration times of the progressive DCGAN and IPD-GAN, as well as five original images randomly selected from the dataset. Figure 6 In the progressive DCGAN iteration-generated images, the clarity is higher than that of the WGAN-GP, and the pitting features are also more obvious. However, it can be seen that compared with the original images, the image damage features are not rich enough, and in the subsequent iterations, the network has common problems in DCGAN training, such as unstable training, loss oscillation, and the generator no longer learning the original image features and generating noise images unrelated to the original image.
[0118] Table 2 Comparative experiments
[0119]
[0120] When the IPD-GAN model iterates 20,000 times, the damaged images generated by the progressive DAGAN model after 20,000 iterations both begin to show the damage characteristics of pitting and flow marks. However, there is still a gap between the learned features and the original images, and the overall image data quality needs to be improved. After 40,000 iterations, compared with the images generated by the progressive DCGAN model, the IPD-GAN model has a more abundant and stable distribution of flow marks and local rust features. To evaluate the image generation effects of the three models more objectively, a comparative experiment was conducted on the generated images after 70,000 iterations of 5,000 rounds of training for the three models, and the FID, PSNR, and IS values of the generated images were calculated respectively, as shown in Table 2. The FID value of the images generated by IPD-GAN is lower than that of the other two models. Compared with the progressive DCGAN, the FID decreased by 31.6, and compared with WGAN-GP, it decreased by 1,424.23, indicating that the feature distribution of the images generated by the IPD-GAN model is closer to the feature distribution of the original images. The PSNR value increased by 0.88 and 2.92 compared with the DCGAN and WGAN-GP models, and the IS value increased by 0.32 and 0.78 respectively, indicating that the IPD-GAN model is not only closer to the original model in terms of feature distribution, but also has richer detailed features in the generated images.
[0121] To verify the effectiveness of the Wasserstein distance loss function with gradient penalty, L1 loss, SSIM loss, and residual blocks in IPD-GAN, ablation experiments were designed. DCGAN was selected as the basic model, and ablation experiments were conducted on each module of this network using the industrial product damage dataset. The FID, PSNR, and IS values of the images generated by different model structures are shown in Table 3. GPW is the Wasserstein distance loss function with gradient penalty, L1-SSIM Loss is the L1 loss and SSIM loss added to the generator, and Resnet is the residual block. Table 3 presents the FID, PSNR, and IS evaluation index values of the DCGAN model, the models with these three improvements added in sequence, and the final model IPD-GAN model of this paper.
[0122] Table 3 Ablation Experiment
[0123]
[0124] As can be seen from Table 3, the calculation results of the three indicators of the images generated by the IPD-GAN model are the best in the ablation experiment. When directly adding the residual blocks in this paper to the DCGAN model, the image quality and diversity of the generated images decline compared with those generated by the DCGAN, indicating that directly adding the residual blocks is not suitable for the original model. Improvement 2 changes the loss function of the DCGAN to the Wasserstein distance loss function with a gradient penalty term. The FID value increases, but the PSNR and IS values also increase. The overall visual quality of the images generated by this model is not as good as that of the DCGAN, but the local differences and the diversity of the images are better than the previous two. The Improvement 3 model adds the L1 loss and the SSIM loss on the basis of the Improvement 2 model. The obtained FID value is improved compared with the previous three models, and the PSNR and IS values are better than the previous two models, but slightly decline compared with the Improvement 3 model. The IPD-GAN model adds residual blocks on the basis of the Improvement 3 model. The results of the three evaluation indicators of the generated image quality are improved compared with the previous four models. Compared with the DCGAN model, the FID value drops from 169.67 to 138.07, a decrease of 18.6%, the PSNR value rises from 15.32 to 16.20, an increase of 5.7%, and the IS value rises from 1.52 to 1.84, an increase of 21.1%. The ablation experiment shows that the overall distribution, local features, and image richness of the images generated by the present invention are more superior.
[0125] Aiming at the problem that the industrial product damage data is difficult to meet the requirements of deep learning training due to the limited sample data of the atmospheric natural environment test, the present invention proposes an IPD-GAN model improved based on the DCGAN to realize the dataset expansion of industrial damage samples. Based on the improvement of the DCGAN model, the loss function is modified to the Wasserstein distance with a gradient penalty term as the adversarial loss, and the L1 loss and the SSIM loss are added to guide the training of the generator, optimizing problems such as unstable training and mode collapse existing in the DCGAN; residual blocks are added to the generator to alleviate the problem of gradient disappearance in the deep network for generating high-resolution images, and at the same time accelerate the training convergence and improve the performance of the generator. The comparative experiment and the ablation experiment show that the overall distribution and richness of the samples generated by the IPD-GAN have been significantly improved, the feature distribution is closer to the real images, and the quality of the generated images is higher, reflecting the reliability of the improved model in this paper for generating industrial product damage data under the atmospheric environment.
[0126] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An industrial product cumulative damage image generation method based on an improved DCGAN, characterized in that, Including the following steps: Step 1: Based on the DCGAN model framework, construct an industrial product cumulative damage image generation network IPD-GAN based on the improved DCGAN. The industrial product cumulative damage image generation network IPD-GAN is provided with a generator and a discriminator; Step 2: The generator obtains random noise, processes the random noise, generates a damaged image, and transmits it to the discriminator; Step 3: The discriminator obtains a real image, discriminates the authenticity of the real image and the damaged image, and uses the Wasserstein distance loss function with a gradient penalty term as the model adversarial loss according to the discrimination result, and combines the L1 loss and the SSIM loss to guide the training of the generator, comprehensively optimizing the system performance; Step 4: Use the trained industrial product cumulative damage image generation network IPD-GAN based on the improved DCGAN to generate damaged images.
2. The industrial product cumulative damage image generation method based on the improved DCGAN according to claim 1, wherein: In the step 1, the generator is provided with a first network layer, a second network layer, a third network layer, a fourth network layer, a fifth network layer, a sixth network layer, a seventh network layer and an eighth network layer connected in sequence; The first network layer, the second network layer, the third network layer, the fourth network layer, the fifth network layer, the sixth network layer and the seventh network layer are all provided with a transposed convolutional layer and a batch normalization layer connected in sequence, and a ReLU activation function is provided after the batch normalization layer; A residual block is also connected after the activation function of the second network layer, the third network layer and the sixth network layer; The eighth network layer is provided with a transposed convolutional layer and a Tanh activation function.
3. The industrial product cumulative damage image generation method based on the improved DCGAN according to claim 2, wherein: The residual block Res is provided with a convolutional layer, a batch normalization layer, a ReLU activation function, a convolutional layer, a batch normalization layer, a residual connection connected in sequence, and outputs after passing through a ReLU activation.
4. A method for generating industrial product cumulative damage images based on an improved DCGAN according to claim 1, characterized in that: In the step 1, the discriminator is provided with a first network block, a second network block, a third network block, a fourth network block, a fifth network block, a sixth network block, a seventh network block and an eighth network block connected in sequence; The first network block is provided with a convolutional layer, and a LeakyReLU activation function is provided after the convolutional layer; The second network block, the third network block, the fourth network block, the fifth network block, the sixth network block and the seventh network block are provided with a convolutional layer and a batch normalization layer connected in sequence, and a LeakyReLU activation function is provided after the batch normalization layer; A dropout layer is connected after the batch normalization layer of the sixth network block and the seventh network block; The eighth network block is provided with a convolutional layer.
5. A method for generating industrial product cumulative damage images based on an improved DCGAN according to claim 1, characterized in that: In the step 3, the expression of the Wasserstein distance loss function is: Among them, W(P r , P g ) is the Wasserstein distance, P r and P g are two probability distributions, Π(P r , P g ) represents the set of all possible joint distribution probabilities of P r and P g ; inf is to take the lower bound; x is the real sample; y is the generated sample; ||x - y|| is the distance between x and y; E represents the expectation, γ is the joint distribution; Introduce a gradient penalty term into the Wasserstein distance loss function. The gradient penalty term L GP has the following expression: where λ is the penalty coefficient; represents the linear interpolation sample between the generated sample and the real sample; is the gradient of the discriminator with respect to the input sample; ||·||2 is the L2 norm.
6. A method for generating industrial product cumulative damage images based on improved DCGAN according to claim 1, characterized in that: In the step 3, the L1 loss L L1 has the following expression: where, ||·||1 is the L1 norm; The SSIM loss L SSIM has the following expression: L SSIM = 1 - SSIM(y, x) SSIM(y,x) = l(x,y)·c(x,y)·s(x,y) where, l(x,y) is the brightness of the image, c(x,y) is the contrast of the image, and s(x,y) is the structure of the image; Among them, μ represents the mean of the sample, σ represents the variance of the sample, and σ xy represents the covariance of the sample; C1 = (K1L) 2 , C2 = (K2L) 2 , C3 = C2 / 2; K1 = 0.01, K2 = 0.03, and L is the pixel value range.
7. A method for generating industrial product cumulative damage images based on an improved DCGAN according to claim 1, characterized in that: The objective formula \(L\) of the loss function of the discriminator D has the following expression: The objective function L of the generator G has the following expression: Among them, represents the expectation of the sample x in the discriminator, represents the expectation of the sample y in the discriminator; λ L1 is the weight coefficient of the L1 loss, and λ SSIM is the weight coefficient of the SSIM loss.
Citation Information
Cited By
Intelligent design method and device for pressure-resistant shell based on full-convolution deep-convolution generative adversarial network fusion process constraint
CN121302541A
Substation fault image generation method and system based on improved DCGAN
CN122530733A