High-resolution bridge crack image generation method based on deep convolutional generative adversarial network

Through deep convolution generation of the deep symmetric architecture and dynamic channel adjustment strategy of the adversarial network, the chessboard effect and detail blurring problems in high-resolution bridge crack image generation are solved, and high-quality bridge crack images are generated, which improves the performance of the detection model.

CN120298981APending Publication Date: 2025-07-11HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510321012.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the generation of high-resolution bridge crack images, the problems of chessboard effect and blurred details, limited feature expression capabilities, excessive consumption of computing resources, insufficient discriminator feature extraction, and difficult to generate high-quality bridge crack images.

Method used

Deep convolution generates adversarial networks, deep symmetric generation adversarial network architecture is designed, combined with dynamic channel adjustment strategies and high-resolution adaptation mechanisms, and replaced transposed convolution through Upsample+Conv2d combination, designed a multi-level feature fusion discriminator, optimized the initial feature size and number of channels, and generated high-quality bridge crack images.

Benefits of technology

The generated bridge crack image details retention capacity is improved by more than 30%, the model convergence speed is improved by 25%, and the parameter quantity is reduced by 20%. The generated images can be used as high-quality training data to solve the problem of scarcity of data in bridge crack detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298981A_ABST
    Figure CN120298981A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision traffic intelligent detection, and particularly discloses a high-resolution bridge crack image generation method based on a deep convolutional generative adversarial network, which combines a deep symmetric network architecture with a dynamic channel adjustment strategy, and effectively balances the calculation efficiency and the feature expression ability. A generator adopts five times of progressive up-sampling (Upsample + Conv2d combination), the chessboard effect of traditional transpose convolution is avoided, and meanwhile, 3 * 3 convolution refining features are introduced after each level of up-sampling, so that the retention capability of a generated 640 * 640 resolution image on high-frequency information such as crack edges and texture details is improved by more than 30%. The structural design is particularly suitable for a bridge crack scene with low contrast and fuzzy edges, and the generated image is closer to the morphological characteristics of a real crack. The method improves the quality and stability of the generated image, expands the bridge crack data, generates different forms of crack images, and solves the problems of scarcity of a crack image data set and the like in bridge crack detection under a complex background.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision traffic intelligent detection, and particularly relates to a method for generating high-resolution bridge crack images based on a deep convolutional generative adversarial network (DCGAN). Background Technique

[0002] Bridge crack detection is a core link for evaluating structural safety. However, the existing technologies face multiple challenges in the field of generating high-resolution crack images. Currently, deep learning-based crack detection models (such as Segformer, TransUNet, etc.) rely heavily on large-scale labeled datasets. However, in actual engineering, bridge crack samples have problems such as sparse distribution (the proportion of crack pixels is less than one-thousandth) and complex backgrounds (affected by lighting, stains, and vegetation occlusion), resulting in the difficulty of the dataset to cover the diversity of real scenarios. Existing solutions mostly use image segmentation or style transfer for data augmentation, but the resolution of the generated images is generally lower than 512*512 pixels, which cannot meet the requirements for calculating the crack width with high precision (the industry standard requires detecting cracks at the 0.1mm level). In addition, traditional data augmentation methods are difficult to simulate the morphological diversity of cracks (such as bifurcation and extension directions) and surface texture details (such as edge serrated features), resulting in the model being prone to false detection and missed detection in actual applications.

[0003] Under this background, the generative adversarial network (GAN) has received attention due to its powerful image generation ability. In order to complete the production of a high-quality bridge crack dataset and meet the deep learning training requirements for bridge crack detection, high-resolution images need to be generated for model training. However, the inherent defects of traditional generative adversarial networks in high-resolution image generation are as follows. First, the first technical bottleneck is the checkerboard effect and detail blurring. The traditional generator uses transposed convolution for upsampling, and its uneven overlapping calculation will cause checkerboard artifacts in the output image, resulting in blurred crack edges and discontinuous textures in the generated image. For example, when using transposed convolution to generate a 640*640 image, high-frequency details (such as crack edges) will appear serrated distortion, and multiple transposed convolution operations will accumulate high-frequency noise, seriously reducing the authenticity of the crack morphology and affecting the training effect of the subsequent segmentation model. Second, the feature expression ability is limited. Existing models mostly adopt symmetric structures with fixed channel numbers (such as DCGAN), which cannot balance the computational efficiency and feature expression ability of high-resolution generation. For example, the computational amount of the generator surges at high-dimensional feature layers (such as 512 channels), while it is difficult to capture the texture features of tiny cracks at low-dimensional layers (such as 32 channels). Third, the training stability is poor. The discriminator is prone to losing crack edge information (such as gradient disappearance) during the multi-layer downsampling process, resulting in an imbalance in the adversarial training between the generator and the discriminator and the occurrence of mode collapse.

[0004] Existing methods still have deficiencies in the high - resolution adaptation mechanism. Currently, mainstream crack generation models (such as SAM, TransUNet) are mainly designed for low - resolution inputs (such as 256*256 pixels). When directly extended to a resolution of 640*640, two major problems will be faced. The first deficiency is the mismatch in potential vector mapping. Traditional methods map potential vectors to initial feature maps of a fixed size (such as 4*4 or 8*8), resulting in loss of details after multiple up - samplings (such as too large a single up - sampling multiple). For example, when directly up - sampling a 4*4 feature map to 640*640, there is a lack of refinement convolution operations in the intermediate layer, and the generated crack texture is rough and has poor continuity. The second deficiency is the excessive consumption of computing resources. High - resolution generation requires a larger batch size and video memory capacity, and existing models do not optimize the dynamic adjustment strategy of the number of channels, resulting in an exponential increase in video memory occupancy (such as the video memory requirement of a 5 - layer network with 512 channels exceeds 24GB).

[0005] Most discriminators of existing GANs use simply stacked convolutional layers for down - sampling. However, for high - resolution images, traditional architectures are difficult to balance the discriminative ability of global semantics and local details. Using convolutional layers with a fixed number of channels in the discriminator results in insufficient feature extraction of high - frequency information such as crack edges in the shallow network, while the deep network loses position sensitivity due to excessive compression of the spatial dimension. Traditional discriminators also do not design a multi - level feature fusion mechanism for the geometric characteristics of cracks (such as directionality, fractal features), making it difficult to accurately measure the difference between the generated image and the real distribution.

[0006] In summary, there are four core problems in the existing technology for high - resolution bridge crack generation: loss of details in high - resolution image generation, checkerboard effect and detail blurring problems in images generated by traditional GANs, insufficient feature extraction ability of discriminators, and the balance problem between feature expression ability and computational efficiency in the high - resolution generation process. Summary of the Invention

[0007] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a method for generating high - resolution bridge crack images based on a deep convolutional generative adversarial network. Based on the generative adversarial network framework, the present invention designs a deep symmetric generative adversarial network architecture, adopts a dynamic channel adjustment strategy and a high - resolution adaptation mechanism to improve the quality and stability of the generated images, expand the bridge crack data, generate crack images of different forms, and solve problems such as the scarcity of crack image datasets in bridge crack detection under complex backgrounds.

[0008] The present invention is realized through the following technical solutions:

[0009] A method for generating high - resolution bridge crack images based on a deep convolutional generative adversarial network includes the following steps:

[0010] Step (1): Design a generative network architecture with deep symmetry. Instead of using transposed convolution for upsampling in the traditional GAN generator, use the combination of Upsample (interpolation) + Conv2d to replace the transposed convolution, avoiding the checkerboard effect. Each layer is enlarged in size through bilinear interpolation (Upsample), and the number of channels is adjusted by 3*3 convolution (Conv2d). The discriminator adopts multi-level feature fusion, designs a 5-layer downsampling module with the number of channels increasing from 32 to 512, and is symmetrically designed with the generator to form an adversarial balance. Each layer realizes downsampling through convolution with a stride of 2, compresses the spatial size, retains the edge gradient information, and retains high-frequency details. Use Dropout2d(0.25) to prevent overfitting and improve the generalization ability.

[0011] Step (2): Dynamic channel adjustment strategy. Adopt the dynamic channel adjustment strategy. At high levels with low resolution, 512 channels are used to encode the global crack trend and branch structure, and at low levels with high resolution, it is reduced to 32 channels, focusing on local texture details and reducing the video memory occupancy.

[0012] Step (3): Design a high-resolution adaptation mechanism, optimize the initial feature size, adopt a high-resolution initial feature map, and through a fully connected layer, map a 100-dimensional latent vector into a 512*20*20 feature tensor (init_size = img_size / / 32) as the input for upsampling.

[0013] Step (4): In order to focus on the crack morphology rather than color information and reduce redundant features, use grayscale images to train the improved deep convolutional generative adversarial network model and perform parameter tuning.

[0014] A further improvement scheme of the present invention is

[0015] The specific operations of Step (1) are as follows:

[0016] Step (1.1): Design a generator sub-network. The generator adopts a 5-layer symmetric upsampling structure, and each layer contains an Upsample + Conv2d(3*3) + BatchNorm + LeakyReLU module. The generator network structure is as Figure 2 shown. The initial input latent vector is mapped into a 512*20*20 feature tensor (corresponding to 1 / 32 size of 640*640 resolution) through a fully connected layer, and then gradually enlarged to 640*640 through 5 times of 2-fold upsampling. Compared with the traditional scheme, only interpolation operation is performed in the upsampling layer, avoiding the problem of uneven weights of transposed convolution; the 3*3 convolution layer refines the features and restores high-frequency details.

[0017] Furthermore, the formulas of the generator and the discriminator are respectively:

[0018]

[0019] Among them, M crack is the crack area mask (enhancing sparsity), λ is the L1 regularization coefficient (default 0.1), which adjusts the weight between the generation quality and the physical constraint. Minimize log(1 - D(G(z))) to make the generated samples be judged as real, and maximize the true data judgment probability log D(x).

[0020] (1.2) Design the discriminator sub-network. The discriminator adopts a 5-layer convolutional downsampling module with a stride of 2, and the number of channels in each layer increases from 32 to 512. The discriminator network structure is as Figure 3 shown. The deep network captures the global semantics (such as the crack trend), and the shallow network retains the edge sharpness. The last layer flattens the features and inputs them into the fully connected layer to output the probability estimate of the image authenticity. During the upsampling feature magnification and refinement process of the generator, the operation formula of each upsampling layer is:

[0021] F l+1 = σ(BN(W l *U(F l )))

[0022] Among them, U(·) is the bilinear upsampling, W l ∈R 3×3×Cin×Cout is the convolutional kernel weight, σ is the LeakyReLU activation function (α = 0.2), and the number of channels decreases according to C l = 512×2 -l decreasing (l = 0, 1,..., 4).

[0023] (1.3) Customize the high resolution for the bridge crack detection requirement. Set the generated resolution as img_size = 640, then dynamically calculate the support, control the latent space dimension through latent_dim = 100, balance the generation diversity and the computational complexity, load the bridge crack images, support grayscale (channels = 1) or RGB format, apply normalization (Normalize([0.5], [0.5])), and uniformly scale the image size to img_size to ensure the input consistency and avoid the model learning deviation caused by the size difference.

[0024] Furthermore, the specific operations of step (2) are as follows:

[0025] In step (2.1), the number of channels of the generator gradually decreases. The number of channels decreases exponentially from 512 to 256, from 256 to 128, from 128 to 64, from 64 to 32, and from 32 to 1. More channels are used in the high-level (low-resolution stage) to encode the global structure, and the number of channels in the low-level (high-resolution stage) is reduced to reduce the computational complexity. At the same time, the local detail generation ability is maintained through small convolutional kernels (3*3). The dynamic channel adjustment strategy formula is as follows. Among them, the generator channel attenuation formula is:

[0026]

[0027] The discriminator channel growth formula is:

[0028]

[0029] (2.2) The number of channels of the discriminator gradually increases. The number of channels starts from 32 and doubles layer by layer. Combining multi-level feature extraction enhances the discrimination accuracy. The optimal channel ratio (doubling each layer) is determined through experiments, so that the gradient information of the crack edge is retained in the shallow layer, and the complex distribution differences are modeled in the deep layer. The discriminator multi-level feature discrimination formula is:

[0030] D k =Dropout(σ(BN(W k*2 D k-1 )))

[0031] Among them, D k-1 represents the output of the previous layer, *2 represents the convolution operation with a stride of 2. The Dropout layer randomly masks neurons to prevent overfitting. First, the gradient is stabilized through the BN layer, and then the σ function is applied to increase the non-linearity. The number of channels increases according to C k =32×2 k (k = 0, 1,..., 4).

[0032] (2.3) High-resolution training is prone to mode collapse and video memory overflow. The number of channels of the generator and the discriminator is adjusted in the opposite direction to balance the computational load. Channels 32-64 are set in the shallow layer to capture high-frequency details such as crack edges and gradient changes, and channels 256-512 are set in the deep layer to model the global semantics of the overall crack distribution and background noise. BatchNorm2d is used to accelerate convergence, Dropout2d is used to prevent overfitting, and LeakyReLU is used to enhance non-linearity. Finally, through the fully connected layer, the 512*20*20 features are compressed into a 1-dimensional true / false probability, and the Sigmoid activation outputs in the range of [0, 1].

[0033] Furthermore, the specific operations in step (3) are as follows:

[0034] (3.1) The latent vector is mapped to a 512*20*20 feature through a fully connected layer. The initial size is 1 / 32 of the 640*640 resolution, rather than the 4*4 or 8*8 initial size in the traditional scheme. After 5 Upsample operations, the 20*20 is gradually enlarged to 640*640. After each level of upsampling, 3*3 convolution is used to refine the features, avoiding the detail blur caused by single upsampling. This design enables the network to encode sufficient spatial information at the early stage, avoiding the loss of details in the subsequent upsampling stage due to too low initial resolution. The formula from the latent vector to the initial feature tensor is:

[0035] F0 = Reshape(W init ·z) ∈ R B×512×20×20

[0036] where z represents the latent space vector, and through the initialized weight matrix W init matrix multiplication is performed. B is the batch size (Batch Size), and the output tensor dimension is B samples × 512 channels × 20×20 spatial resolution.

[0037] (3.2) Adjust the activation function. LeakyReLU enhances the nonlinearity, and Tanh limits the output to [-1,1], matching the normalized data. The generator network passes through 5 upsampling modules, each layer containing Upsample + Conv2d + BatchNorm + LeakyReLU, and the number of channels decreases from 512 to 32, balancing the computational efficiency and feature expression ability.

[0038] Furthermore, the specific operations of parameter tuning in step (4) are as follows:

[0039] Set epochs to 200, batch_size to 16, the number of cpu threads to 8, latent_dim to 100, and the most crucial img_size to 640. In order to focus on the crack morphology rather than color information and reduce redundant features, grayscale images are used, and channels are set to 1.

[0040] The beneficial effects of the present invention are:

[0041] Compare the performance of the present invention with that of existing mainstream GAN models in the bridge crack generation task, and the results are summarized as Figure 1 .

[0042] Aiming at the technical bottlenecks existing in the existing generative adversarial networks in the generation of high-resolution bridge crack images, the present invention provides a deep symmetric network architecture and an optimization method capable of efficiently generating high-quality and high-fidelity bridge crack images. The present invention designs a deep symmetric network architecture, sets up a 5-layer symmetric generative adversarial network architecture, and the generator and discriminator adopt a 5-layer symmetric structure. The combination of upsampling (Upsample) and standard convolution is used to replace the transposed convolution, avoiding edge blurring caused by frequency domain aliasing and eliminating the checkerboard effect of the transposed convolution. A discriminator structure with increasing number of channels is designed. The discriminator gradually compresses features through 5 times of downsampling (stride 2 convolution), retains high-frequency details such as crack edges, and enhances the multi-level feature extraction ability of the discriminator. A dynamic channel adjustment strategy is adopted. The number of channels of the generator decreases from 512 to 32 to balance the computational efficiency and feature expression ability; the number of channels of the discriminator increases from 32 to 512, and the discriminant ability is enhanced through multi-level feature extraction. Adaptive matching of the number of channels and resolution is achieved in the generator and discriminator, optimizing the allocation of computing resources and improving the training efficiency. A high-resolution adaptation mechanism is adopted. The latent vector is mapped into a 512*20*20 feature through a fully connected layer, with an initial size of 1 / 32 of the 640*640 resolution, matching the high-resolution image generation method. After 5 times of Upsample operations, 20*20 is gradually enlarged to 640*640, and a 3*3 convolution is used to refine the features after each upsampling to avoid detail blurring caused by a single upsampling. Through the high-resolution initial feature mapping (512*20*20) and progressive upsampling, the generation of the microscopic morphology of the crack is accurately controlled.

[0043] The high-resolution bridge crack image generation method based on the deep convolutional generative adversarial network proposed by the present invention, the combination of the deep symmetric network architecture and the dynamic channel adjustment strategy effectively balances the computational efficiency and the feature expression ability. The generator adopts 5 times of progressive upsampling (Upsample+Conv2d combination), avoiding the checkerboard effect of the traditional transposed convolution. At the same time, a 3*3 convolution is introduced to refine the features after each upsampling, so that the ability to retain high-frequency information such as crack edges and texture details in the generated 640*640 resolution image is increased by more than 30%. This structural design is particularly suitable for bridge crack scenarios with low contrast and blurred edges, and the generated images are closer to the morphological characteristics of real cracks. The discriminator adopts a 5-layer downsampling structure with an increasing number of channels from 32 to 512, and the discriminant ability is enhanced through multi-level feature extraction. Compared with the traditional single-stage discriminant network, this design improves the model convergence speed by about 25%, and the dynamic channel adjustment strategy (the number of channels of the generator decreases and the discriminator increases) reduces the model parameter quantity by 20%. The images generated by this method can be used as a supplement to high-quality training data, and when the sample size is insufficient, the detection model can obtain a large number of realistic data sets for testing. Combining the data augmentation method of the style transfer network can construct a large-scale data set containing complex backgrounds and solve the problem of difficult crack data acquisition in practical engineering. Brief Description of the Drawings

[0044] Figure 1 This is the comparison result of the technical solution of the present invention and the performance of existing mainstream GAN models in the task of bridge crack generation;

[0045] Figure 2 This is the structural diagram of the generator network in the method of the present invention;

[0046] Figure 3 This is the structural diagram of the discriminator network in the method of the present invention;

[0047] Figure 4 This is the training result of the 1st round in the embodiment of the present invention;

[0048] Figure 5 This is the training result of the 100th round in the embodiment of the present invention;

[0049] Figure 6 This is the training result of the 200th round in the embodiment of the present invention;

[0050] Figure 7 This is the 64 - resolution image of the unimproved model in the embodiment of the present invention;

[0051] Figure 8 This is the 640 - resolution image after improvement in the embodiment of the present invention. Detailed Description of the Invention

[0052] The present invention will be introduced in detail below in conjunction with specific embodiments.

[0053] Embodiment 1: A method for generating high - resolution bridge crack images based on a deep convolutional generative adversarial network

[0054] Step (1) Design a deep symmetric generator network architecture. Instead of using transposed convolution for upsampling in the traditional GAN generator, use the Upsample (interpolation) + Conv2d combination to replace the transposed convolution to avoid the checkerboard effect. Each layer is enlarged in size through bilinear interpolation (Upsample), and the number of channels is adjusted by 3 * 3 convolution (Conv2d); the discriminator adopts multi - level feature fusion, and designs 5 downsampling modules with the number of channels increasing from 32 to 512. It is symmetrically designed with the generator to form an adversarial balance. Each layer realizes downsampling through convolution with a stride of 2, compresses the spatial size, retains the edge gradient information, and retains high - frequency details. Use Dropout2d(0.25) to prevent overfitting and improve the generalization ability.

[0055] Including step (1.1) Design the generator sub - network. The generator adopts a 5 - layer symmetric upsampling structure, and each layer contains an Upsample + Conv2d(3 * 3) + BatchNorm + LeakyReLU module. The generator network structure is asFigure 2 As shown in the figure. The initial input latent vector is mapped to a 512×20×20 feature tensor (corresponding to 1 / 32 of the 640×640 resolution) through a fully connected layer, and then gradually enlarged to 640×640 through 5 times of 2x upsampling. Compared with the traditional scheme, the upsampling layer only performs interpolation operations to avoid the problem of uneven weights in transposed convolution; the 3×3 convolution layer refines the features and restores high-frequency details. The formulas of the generator and discriminator are as follows:

[0056]

[0057] where M crack is the crack area mask (enhancing sparsity), λ is the L1 regularization coefficient (default 0.1), which adjusts the weights between generation quality and physical constraints. Minimizing log(1 - D(G(z))) makes the generated samples be judged as real, and maximizing the true data judgment probability log D(x).

[0058] Step (1.2) designs the discriminator sub-network. The discriminator adopts a 5-layer stride-2 convolution downsampling module, and the number of channels in each layer increases from 32 to 512. The discriminator network structure is as Figure 3 shown in the figure. The deep network captures global semantics (such as the crack trend), and the shallow network retains edge sharpness. The last layer flattens the features and inputs them into a fully connected layer to output the probability estimate of the image authenticity. During the process of upsampling and refining the features of the generator, the operation formula of each upsampling layer is:

[0059] F l+1 = σ(BN(W l *U(F l )))

[0060] where U(·) is bilinear upsampling, W l ∈R 3×3×Cin×Cout is the convolution kernel weight, σ is the LeakyReLU activation function (α = 0.2), and the number of channels decreases according to C l = 512×2 -l (l = 0, 1, …, 4).

[0061] Step (1.3) customizes the high resolution for bridge crack detection. Set the generated resolution img_size = 640, then dynamically calculate the support, control the latent space dimension through latent_dim = 100, balance the generation diversity and computational complexity, load the bridge crack images, support grayscale (channels = 1) or RGB format, apply normalization (Normalize([0.5], [0.5])), and uniformly scale the image size to img_size to ensure input consistency and avoid model learning bias caused by size differences.

[0062] Step (2) Dynamic channel adjustment strategy. Adopt the dynamic channel adjustment strategy. For the low-resolution of the high layer, 512 channels are used to encode the global crack orientation and branch structure. For the high-resolution of the low layer, the number of channels is reduced to 32 to focus on local texture details and reduce video memory occupancy.

[0063] It includes step (2.1) The channels of the generator gradually decrease. The number of channels decreases exponentially according to the rule of 512 to 256, 256 to 128, 128 to 64, 64 to 32, 32 to 1. More channels are used in the high layer (low-resolution stage) to encode the global structure, and the number of channels in the low layer (high-resolution stage) is reduced to reduce the computational complexity. At the same time, the local detail generation ability is maintained through small convolutional kernels (3*3). The formula for the dynamic channel adjustment strategy is as follows. Among them, the formula for the decay of the generator channels is:

[0064]

[0065] The formula for the growth of the discriminator channels is:

[0066]

[0067] Step (2.2) The channels of the discriminator gradually increase. The number of channels starts from 32 and doubles layer by layer to enhance the discrimination accuracy by combining multi-level feature extraction. The optimal channel ratio (doubling each layer) is determined through experiments to retain the gradient information of the crack edge in the shallow layer and model the complex distribution differences in the deep layer. The formula for the multi-level feature discrimination of the discriminator is:

[0068] D k =Dropout(σ(BN(W k*2 D k-1 )))

[0069] Among them, D k-1 represents the output of the previous layer, *2 represents the convolution operation with a stride of 2. The Dropout layer randomly masks neurons to prevent overfitting. First, the gradient is stabilized through the BN layer, and then the σ function is applied to increase the non-linearity. The number of channels increases according to C k =32×2 k (k = 0, 1, …, 4).

[0070] In step (2.3), mode collapse and video memory overflow are likely to occur during high-resolution training. The channel numbers of the generator and the discriminator are adjusted in the opposite direction to balance the computational load. 32-64 channels are set in the shallow layer to capture high-frequency details such as crack edges and gradient changes, and 256-512 channels are set in the deep layer to model the global semantics of the overall crack distribution and background noise. BatchNorm2d is used to accelerate convergence, Dropout2d is used to prevent overfitting, and LeakyReLU is used to enhance non-linearity. Finally, through the fully connected layer, the 512*20*20 features are compressed into a 1-dimensional true / false probability, and the Sigmoid activation outputs in the range of [0, 1].

[0071] Step (3) Design a high-resolution adaptation mechanism to optimize the initial feature size. Adopt a high-resolution initial feature map. Through the fully connected layer, map the 100-dimensional latent vector into a 512*20*20 feature tensor (init_size = img_size / / 32), which is used as the input for upsampling.

[0072] It includes step (3.1)

[0073] The latent vector is mapped into a 512*20*20 feature through the fully connected layer. The initial size is 1 / 32 of the 640*640 resolution, rather than the 4*4 or 8*8 initial size in the traditional scheme. After 5 Upsample operations, the 20*20 is gradually enlarged to 640*640. After each level of upsampling, a 3*3 convolution is used to refine the features to avoid detail blurring caused by single upsampling. This design enables the network to encode sufficient spatial information at an early stage, avoiding detail loss in the subsequent upsampling stage due to too low initial resolution. The formula from the latent vector to the initial feature tensor is:

[0074] F0 = Reshape(W init ·z) ∈ R B×512×20×20

[0075] where z represents the latent space vector, and through the initialized weight matrix W init perform matrix multiplication. B is the batch size (Batch Size), and the output tensor dimension is B samples × 512 channels × 20 × 20 spatial resolution.

[0076] Step (3.2) Adjust the activation function. LeakyReLU enhances the non-linearity, and Tanh restricts the output to [-1,1] to match the normalized data. The generator network passes through 5 upsampling modules, each layer containing Upsample+Conv2d+BatchNorm+LeakyReLU, and the number of channels decreases from 512 to 32 to balance the computational efficiency and feature expression ability.

[0077] Step (4) To focus on the crack morphology rather than color information and reduce redundant features, use grayscale images to train the improved deep convolutional generative adversarial network model and perform parameter tuning.

[0078] Include step (4.1) to set epochs to 200, batch_size to 16, the number of CPU threads to 8, latent_dim to 100, and the most crucial img_size to 640. To focus on crack morphology rather than color information and reduce redundant features, grayscale images are used, and channels are set to 1. Store 1000 real crack image datasets in the data_crack folder, and the model reads the image information in the folder to start training. The training result of the 10th round is as Figure 4 shown. The image is blurry and not yet formed. The training result of the 100th round is as Figure 5 shown. The edges are relatively blurry, but the general outline is already clear. The training result after the 200th round is as Figure 6 shown, which is very close to the real crack image. Use the.pth weight file generated by the trained model for image generation. Select several generated images and compare them with the images trained by the generative adversarial network model with a resolution of 64*64. The generated 64-pixel image is as Figure 7 shown, and the generated 640-pixel image is as Figure 8 shown. The images generated by the high-resolution 640*640 model are closer to the real images in terms of both clarity and vividness.

[0079] In summary, the present invention first adopts a symmetric network architecture and a progressive upsampling mechanism, sets up a 5-layer symmetric generative adversarial network architecture, the generator and the discriminator adopt a 5-layer symmetric structure, and the generator realizes the resolution doubling through 5 times of upsampling (the combination of Upsample+Conv2d), avoiding the checkerboard effect of transposed convolution; the discriminator gradually compresses the features through 5 times of downsampling (stride 2 convolution), retains high-frequency details such as crack edges, and forms a closed-loop encoding-decoding structure, significantly improving the stability of adversarial training and realizing detail fidelity at high resolution. Secondly, the dynamic channel number adjustment strategy takes into account both computational efficiency and feature expression ability. The number of channels of the generator decreases from 512 to 32 to balance computational efficiency and feature expression ability; the number of channels of the discriminator increases from 32 to 512 to enhance the discrimination ability through multi-level feature extraction, breaking through the limitations of traditional fixed channel design. Finally, a loss function and an initial feature size are designed for the sparse characteristics of bridge cracks to improve the physical rationality of the generated images. In the bridge crack detection scenario, the method generates 640*640 resolution images, and the detail retention ability is better than that of the Pix2PixHD model. The generated crack images can be directly used for data augmentation, increasing the mAP of detection models such as YOLOv5 by 9%.

[0080] The above embodiments are only for explaining the technical concept and features of the present invention, and the purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention shall be covered within the protection scope of the present invention.

Claims

1. A high-resolution bridge crack image generation method based on a deep convolutional generative adversarial network, characterized in that, It includes the following steps: Step (1): Design a depth-symmetric generation network architecture. The generator uses the Upsample + Conv2d combination for upsampling. Each layer is enlarged in size through bilinear interpolation, and the number of channels is adjusted by 3*3 convolution. The discriminator uses multi-level feature fusion and designs a 5-layer downsampling module with the number of channels in each layer increasing from 32 to 512. It is symmetrically designed with the generator to form an adversarial balance. Each layer realizes downsampling through convolution with a stride of 2, compresses the spatial size, retains edge gradient information, retains high-frequency details, and uses Dropout2d(0.25) to prevent overfitting and improve generalization ability; Step (2): Dynamic channel adjustment strategy. Adopt the dynamic channel adjustment strategy. At high levels with low resolution, 512 channels are used to encode the global crack trend and branch structure. At low levels with high resolution, it is reduced to 32 channels to focus on local texture details and reduce video memory occupancy; Step (3): Design a high-resolution adaptation mechanism, optimize the initial feature size, adopt a high-resolution initial feature map, and map a 100-dimensional latent vector to a 512*20*20 feature tensor (init_size = img_size / / 32) through a fully connected layer as the input for upsampling; Step (4): Use grayscale images to train the improved deep convolutional generative adversarial network model and perform parameter tuning.

2. The high-resolution bridge crack image generation method based on a deep convolutional generative adversarial network according to claim 1, characterized in that: Specifically, step (1) is as follows: (1.1) Design the generator sub-network. The generator adopts a 5-layer symmetric upsampling structure. Each layer contains the Upsample + Conv2d(3*3) + BatchNorm + LeakyReLU module. The initial input latent vector is mapped to a 512*20*20 feature tensor through a fully connected layer, and then gradually enlarged to 640*640 through 5 times of 2-fold upsampling; the 3*3 convolutional layer refines the features and restores high-frequency details; (1.2) Design the discriminator sub-network. The discriminator adopts a 5-layer convolution downsampling module with a stride of 2. The number of channels in each layer is 512. The deep network captures global semantics, and the shallow network retains edge sharpness. The features are flattened and input into the fully connected layer in the last layer to output the probability estimate of the image authenticity; (1.3) Customize high resolution for the bridge crack detection requirement. Set the generated resolution to img_size = 640, then dynamically calculate the support, control the latent space dimension through latent_dim = 100, balance the generation diversity and computational complexity, load the bridge crack images, support grayscale or RGB formats, apply normalization (Normalize([0.5],[0.5])), and uniformly scale the image size to img_size to ensure input consistency and avoid model learning deviation caused by size differences.

3. The high-resolution bridge crack image generation method based on a deep convolutional generative adversarial network according to claim 2, wherein: The formula of the generator is: The formula of the generator is: Among them, M crack is the crack area mask (enhancing sparsity), λ is the L1 regularization coefficient (default 0.1), which adjusts the weights between the generation quality and physical constraints. Minimize log(1 - D(G(z))) to make the generated samples be judged as real, and maximize the true data judgment probability log D(x).

4. The high-resolution bridge crack image generation method based on a deep convolutional generative adversarial network according to claim 2, characterized in that: During the process of upsampling feature amplification and refinement of the generator, the operation formula of each upsampling layer is: F l+1 = σ(BN(W l *U(F l ))) Among them, U(·) is bilinear upsampling, and W l ∈R 3×3×Cin×Cout is the convolutional kernel weight, σ is the LeakyReLU activation function (α = 0.2), and the number of channels decreases according to C l = 512 × 2 -l decreasingly (l = 0, 1, …, 4).

5. The high-resolution bridge crack image generation method based on a deep convolutional generative adversarial network according to claim 1, characterized in that: Specifically, step (2) is as follows: (2.1) The generator channels gradually decrease. The number of channels decreases exponentially from 512 to 256, from 256 to 128, from 128 to 64, from 64 to 32, and from 32 to 1. More channels are used in the high layers to encode the global structure, and the number of channels in the low layers is reduced to lower the computational complexity. At the same time, the local detail generation ability is maintained through small convolution kernels. (2.2) The discriminator channels gradually increase. The number of channels doubles layer by layer starting from 32. By combining multi-level feature extraction, the discrimination accuracy is enhanced. The optimal channel ratio is determined through experiments to retain the gradient information of the crack edge in the shallow layer and model the complex distribution differences in the deep layer. (2.3) High-resolution training is prone to mode collapse and out-of-memory errors. The number of channels of the generator and the discriminator is adjusted in the opposite direction to balance the computational load. Channels 32 - 64 are set in the shallow layer to capture high-frequency details such as crack edges and gradient changes, and channels 256 - 512 are set in the deep layer to model the global semantics of the overall crack distribution and background noise. BatchNorm2d is used to accelerate convergence, Dropout2d is used to prevent overfitting, and LeakyReLU is used to enhance non-linearity. Finally, through the fully connected layer, the 512*20*20 features are compressed into a 1-dimensional true / false probability, and Sigmoid activation outputs in the range of [0,1].

6. The high-resolution bridge crack image generation method based on a deep convolutional generative adversarial network according to claim 5, wherein: The formula for the decay of the generator channels is: The formula for the growth of the discriminator channels is:

7. The high-resolution bridge crack image generation method based on a deep convolutional generative adversarial network according to claim 5, wherein: The formula for the multi-level feature discrimination of the discriminator is: D k = Dropout(σ(BN(W k * 2D k-1 ))) Among them, D k-1 represents the output of the previous layer. *2 represents the convolution operation with a stride of 2. The Dropout layer randomly masks neurons to prevent overfitting. First, the gradient is stabilized through the BN layer, and then the σ function is applied to increase non-linearity. The number of channels increases according to C k = 32 × 2 k incrementally (k = 0, 1, …, 4).

8. The high-resolution bridge crack image generation method based on a deep convolutional generative adversarial network according to claim 1, characterized in that: Step (3) specifically is: (3.1) The latent vector is mapped to 512*20*20 features through the fully connected layer. The initial size is 1 / 32 of the 640*640 resolution, rather than the traditional 4*4 or 8*8 initial size. After 5 Upsample operations, the 20*20 is gradually enlarged to 640*640, and a 3*3 convolution is used to refine the features after each upsampling. (3.2) Adjust the activation function. LeakyReLU enhances non-linearity, and Tanh restricts the output to [-1,1] to match the normalized data. The generator network passes through 5 upsampling modules, each containing Upsample+Conv2d+BatchNorm+LeakyReLU, and the number of channels decreases from 512 to 32 to balance the computational efficiency and feature expression ability.

9. The high-resolution bridge crack image generation method based on a deep convolutional generative adversarial network according to claim 8, characterized in that: The formula from the latent vector to the initial feature tensor is: F0 = Reshape(W init ·z) ∈ R B×512×20×20 where z represents the latent space vector, and through initializing the weight matrix W init matrix multiplication is performed, B is the batch size, and the output tensor dimension is B samples × 512 channels × 20 × 20 spatial resolution.

10. The high-resolution bridge crack image generation method based on a deep convolutional generative adversarial network according to claim 1, wherein: In step (4), set epochs to 200, batch_size to 16, the number of CPU threads to 8, latent_dim to 100, and the most crucial img_size to 640. To focus on the crack morphology rather than color information and reduce redundant features, grayscale images are used, and channels are set to 1.