A bridge steel structure disease visual image data processing method

CN122737720APending Publication Date: 2026-09-11HUBEI HUICHUANG HEAVY ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610778652.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0002]当前基于机器视觉的桥梁钢结构病害识别技术,其标注数据通常根据某一类病害数据进行建立,数据集中于单一损伤类型,以及基于实验环境或者预设状态条件的理想环境来模拟生成,难以实现对实际多病害类型以及复杂场景下的有效识别,难以实现在实际应用场景的快速落地实施,在应用过程中更需要进行长期的频繁地本地化针对性优化调试,导致工作量极其庞大,但最终识别分析效果并不好;同时,实际现场环境下多变的光照要素、环境要素、表面状态、遮挡以及各类干扰要素的存在,导致图像状态复杂多变,在进行图像处理过程中,直接应用基于实验数据声场和验证的算法难以应满足实际分析需求

Benefits of technology

[0003]本发明的目的在于,基于实际需求。提供一种基于对现有桥梁钢结构病害图像数据进行扩增优化处理后基于多尺度空间特征进行分割识别,以实现对桥梁钢结构进行病害区域识别和分割的方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122737720A_ABST
    Figure CN122737720A_ABST
Patent Text Reader

Abstract

This invention belongs to the technical field of bridge steel structure defect analysis and processing, and particularly relates to a method for processing visual image data of bridge steel structure defects. The method includes the following steps: using a generative adversarial network to learn from an existing visual image dataset of bridge steel structure defects, expanding and generating a new dataset; completing the bridge defect image classification label based on prior knowledge; constructing a multi-scale segmentation network output mapping map, consisting of a symmetrically configured encoding network, a decoding network, a connection network located within the encoding and decoding networks, and a synthesis network; and finally using the OTSU algorithm to complete the binarization segmentation. This application aims to achieve rapid brightening, classification, and defect region segmentation extraction of bridge defect images acquired in real-time with low capacity and low information density, enabling better implementation and meeting the analysis needs of bridge steel structure defects in engineering projects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bridge steel structure defect analysis and processing technology, and in particular relates to a method for processing visual image data of bridge steel structure defects. Background Technology

[0002] Current machine vision-based bridge steel structure defect identification technologies typically use labeled data built based on a specific type of defect. This data is concentrated on a single damage type and is generated through simulations of ideal environments based on experimental conditions or preset state conditions. This makes it difficult to effectively identify multiple defect types and complex scenarios in reality, hindering rapid implementation in practical applications. Furthermore, the application process requires long-term, frequent, and targeted local optimization and debugging, resulting in an extremely large workload, yet the final identification and analysis results are still unsatisfactory. Simultaneously, the presence of variable lighting, environmental factors, surface conditions, occlusion, and various interference factors in actual field environments leads to complex and variable image states. Directly applying algorithms based on experimental data and validation during image processing is insufficient to meet actual analysis needs. Summary of the Invention

[0003] The purpose of this invention is to provide a method for identifying and segmenting defect areas in bridge steel structures based on multi-scale spatial features after augmenting and optimizing existing bridge steel structure defect image data, thereby meeting practical needs.

[0004] To achieve the above objectives, the present invention adopts the following technical solution.

[0005] A method for processing visual image data of defects in bridge steel structures includes the following steps:

[0006] Step 1: Enhancement and optimization of bridge steel structure defect image data. Generative adversarial network is used to learn from the existing bridge steel structure defect visual image dataset and expand it to generate a bridge steel structure defect dataset.

[0007] Step 2: Multi-label image classification of bridge steel structure defects by integrating prior knowledge

[0008] Based on prior knowledge, image classification and labeling of bridge steel structure defects are completed. A label feature learning network composed of two-layer graph convolutional networks is used to learn the latent semantic associations and co-occurrence rules between different defect labels, and output the semantic feature vector of label relationship. At the same time, a multi-axis self-attention network based on MaxViT is used to capture global contextual information in the image and output the high-dimensional feature vector of the image. The visual feature vector extracted from the image and the semantic feature vector learned from the label relationship are fed into a multimodal feature fusion module for deep fusion, and then the classification result is output by a classifier, and the result corresponds to the probability of the defect category of the image.

[0009] Step 3: Image segmentation of bridge steel structure defects

[0010] Construct a multi-scale segmentation network consisting of symmetrically configured encoding networks, decoding networks, connection networks located in the encoding and decoding networks, and a synthesis network;

[0011] The encoder network obtains multi-scale disease features by compressing layer by layer through convolutional and pooling layers to capture contextual information, while the decoder network restores the feature scale layer by layer through upsampling and convolutional layers to accurately locate the position of each pixel. The synthesis network extracts the fusion feature output of the decoding unit and uses the spatial attention module to extract the spatial fusion features. The two output features are mapped to a probability map, and finally the OTSU algorithm is used to complete the binarization segmentation.

[0012] A further improvement or preferred embodiment of the aforementioned method for processing visual image data of bridge steel structure defects, specifically step one, enhancing and optimizing the image data of bridge steel structure defects, includes:

[0013] a. Collect images of bridge steel structure defects and perform necessary preprocessing, including grayscale unification and filtering for noise reduction; establish an original dataset of bridge steel structure defect images. ;

[0014] b. Augmenting the bridge steel structure defect images in the original dataset of beam steel structure defect images using DCGAN; the DCGAN adds a convolutional neural network to the adversarial generative network based on the traditional generator and discriminator; the loss function of the generative network is designed as follows: Where n is the total amount of data in the dataset; This refers to pseudo-images generated by a generative network; This represents the probability that the current image is identified as a fake image. For random noise; design the loss function of the discriminant network as follows: ;

[0015] c. Merge the generated images with the original images to construct a dataset of bridge steel structure defects.

[0016] In a further improvement or preferred embodiment of the aforementioned method for processing visual image data of bridge steel structure defects, in step a, when the original data comes from different acquisition terminals and configuration parameters, adaptive cropping and scaling of all images are performed according to the preset image aspect ratio, while unifying the range of image pixel values ​​to ensure the consistency of the original data.

[0017] In a further improved or preferred embodiment of the aforementioned method for processing visual image data of bridge steel structure defects, in step three, bridge steel structure defect image segmentation,

[0018] The coding network consists of m+1 ascending connected coding units. The decoding network consists of m decoding units that are symmetrically arranged with respect to the first m encoding units and connected in reverse order. The network connection includes setting the kth ( m connection units between the m-th encoding unit and the (m-k+1)th decoding unit ;

[0019] The encoding unit consists of a 3x3 convolutional layer and a 2x2 pooling layer with a stride of 2, where the i-th ( ) coding units The input is the (i-1)th encoding module. The output feature map is obtained by downsampling by 0.5 times; the decoding unit packet consists of a 2x2 upsampling layer with a stride of 2 and two 3x3 convolutional layers; the connection unit consists of two 3x3 convolutional layers.

[0020] The first decoding unit is obtained by concatenating the outputs of the m-th and (m+1)-th encoding units and then extracting and fusing them through the connection unit.

[0021] The kth ( The input of the (k-1)th decoding unit is the output of the (k-1)th decoding unit and the output of the (k)th encoding unit, which are then concatenated and extracted and fused by the connection unit.

[0022] A further improvement or preferred embodiment of the aforementioned method for processing visual image data of bridge steel structure defects involves a synthetic network acquiring the fused feature output of the m-th decoding unit. and output the fused features The spatial fusion feature output is obtained by using the spatial attention module as input. The two fused feature outputs are finally processed by convolution compression and normalization and mapped to a channel with the same number of disease types to generate a defect prediction probability mapping map. Finally, the OTSU algorithm is used to complete the binarization segmentation. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the principle and structure of the CGAN deep convolutional generative adversarial network;

[0024] Figure 2 This is a schematic diagram of a multi-scale segmentation and extraction network structure;

[0025] Figure 3 These are images of surface damage to the steel structure of a bridge.

[0026] Figure 4 This is the result of segmentation and identification of surface damage on the bridge steel structure. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0028] This application relates to a visual image data processing method for bridge steel structure defects, which is mainly used to achieve rapid brightening, classification, and image defect region segmentation and extraction of bridge defect images acquired in real time with low capacity and low information density, so as to better complete the implementation and meet the analysis needs of bridge steel structure defects in engineering.

[0029] The following detailed explanation of technical solutions not applicable to this application is provided in conjunction with specific process steps.

[0030] Step 1: Enhancement and optimization of image data of bridge steel structure defects

[0031] Due to differences in specific bridge structures and operating environments, the distribution types and progression of defects vary significantly among different types of bridge steel structures. Therefore, classification models trained on traditional, general-purpose large datasets often suffer from complexity, numerous parameters, poor specificity, and insufficient adaptability. To achieve targeted analysis and processing based on real-time acquired images of bridge steel structure defects, while ensuring the effectiveness of low-sample-count image data in subsequent analysis, this application utilizes Generative Adversarial Networks (DCGANs) to learn from existing small datasets, expanding and generating new, high-quality images to improve the accuracy and robustness of subsequent model training. Specific steps include:

[0032] a. Collect images of bridge steel structure defects and perform necessary preprocessing, including grayscale unification and filtering for noise reduction; establish an original dataset of bridge steel structure defect images. ;

[0033] When the raw data comes from different acquisition terminals and configuration parameters, all images should be adaptively cropped and scaled according to the preset image aspect ratio, while unifying the range of image pixel values ​​to ensure the consistency of the raw data.

[0034] b. Augmentation of bridge steel structure defect images in the original dataset of beam steel structure defect images using DCGAN; the DCGAN adds a convolutional neural network to the adversarial generative network based on the traditional generator and discriminator.

[0035] The generator in the DCGAN deep convolutional adversarial generative network processes a random noise image through a deconvolutional neural network to generate an image with the same size as the original image of the bridge steel structure. This generated image is then mixed with the original image data and classified by a discriminant network composed of continuous convolutions. The similarity of features between the generated and original images is determined, and the parameters of the DCGAN deep convolutional adversarial generative network are optimized based on changes in the loss function. The basic structure of the DCGAN deep convolutional adversarial generative network is as follows: Figure 1 As shown.

[0036] To measure the error in the adversarial training process of DCGAN and backpropagate the error to update the parameters in the network, the errors of the loss function generator and discriminator are constructed separately. Specifically:

[0037] Considering that the goal of the DCGAN adversarial training network's generative network is to generate images with feature distributions similar to real images, and to minimize the probability of generating fake images, the loss function of the generative network is designed as follows: Where n is the total amount of data in the dataset; This refers to pseudo-images generated by a generative network; This represents the probability that the current image is identified as a fake image. It is random noise;

[0038] Meanwhile, the goal of the discriminative network in the DCGAN adversarial training network is to accurately distinguish the true and false attributes of the input image and improve the accuracy of the judgment. Therefore, the loss function of the discriminative network is designed as follows: ;

[0039] c. Merge the generated images with the original images to construct a dataset of bridge steel structure defects;

[0040] To achieve the classification and recognition of bridge defect image types, this application adopts a multi-label bridge steel structure defect image classification method that integrates prior knowledge to analyze and automatically classify image defects in the constructed bridge steel structure defect dataset. Specifically:

[0041] Step 2: Multi-label image classification of bridge steel structure defects by integrating prior knowledge

[0042] Based on prior knowledge, we completed the image classification and labeling of bridge defects in the bridge steel structure defect dataset. We used a label feature learning network composed of two-layer graph convolutional networks to learn the latent semantic associations and co-occurrence rules between different defect labels and output the semantic feature vector of label relationship. At the same time, we used a multi-axis self-attention network based on MaxViT to capture global contextual information in the image and output the high-dimensional feature vector of the image.

[0043] The visual feature vectors extracted from the image and the semantic feature vectors learned from the label relationships are fed into a multimodal feature fusion module for deep fusion. The classifier then outputs the classification result, which corresponds to the probability of the disease category in the image, thus achieving multi-label classification. Generally, the classifier is built using a fully connected layer with a sigmoid activation function.

[0044] Existing image recognition and analysis methods include whole-image classification, object detection networks, and pixel segmentation networks. Among them, whole-image classification has high requirements for image quality and the feature recognition accuracy of the target object in the image, making it difficult to complete the task of identifying and extracting detailed features of steel structure defects. The object detection network-based method has poor recognition effect on early-stage defects such as small cracks and rusted areas, and there is a problem of confusing regional defects with surface features of the steel structure. Therefore, in order to meet the requirements of more accurate and effective detection and analysis of bridge steel structure defects throughout the entire life cycle, this application preferably adopts a deep learning network based on multi-scale segmentation network for pixel-by-pixel recognition. The specific structure and principle are described in step three.

[0045] Step 3: Pixel-level image segmentation of bridge steel structure defects

[0046] A multi-scale segmentation network is constructed. The multi-scale segmentation network is based on a symmetric encoding and decoding structure, consisting of a symmetrically set encoding network, a decoding network, a connection network located between the encoding network and the decoding network, and a synthesis network.

[0047] The encoder network obtains multi-scale disease features through layer-by-layer compression to capture contextual information, while the decoder network gradually recovers the feature scale to accurately locate the position of each pixel, and the synthesis network extracts multi-scale features to complete the output; the encoding network includes m+1 sequentially connected encoding units. The decoding network consists of m decoding units that are symmetrically arranged with respect to the first m encoding units and connected in reverse order. The network connection includes setting the kth ( m connection units between the m-th encoding unit and the (m-k+1)th decoding unit ;

[0048] The encoding unit consists of a 3x3 convolutional layer and a 2x2 pooling layer with a stride of 2, where the i-th ( ) coding units The input is the (i-1)th encoding module. The output feature map is obtained by downsampling by 0.5 times; the decoding unit packet consists of a 2x2 upsampling layer with a stride of 2 and two 3x3 convolutional layers; the connection unit consists of two 3x3 convolutional layers.

[0049] The first decoding unit is obtained by concatenating the outputs of the m-th and (m+1)-th encoding units and then extracting and fusing them through the connection unit.

[0050] The kth ( The input of the (k-1)th decoding unit is the output of the (k-1)th decoding unit and the output of the kth encoding unit, which are then concatenated and extracted and fused by the connection unit to obtain the result.

[0051] The synthesis network obtains the fused feature output of the m-th decoding unit. and output the fused features The spatial fusion feature output is obtained by using the spatial attention module as input. The two fused feature outputs are then processed through convolutional compression and normalization, mapped to channels corresponding to the number of disease types to generate a defect prediction probability mapping. Finally, the OTSU algorithm is used to complete binarization segmentation, achieving region segmentation and extraction. Figure 3 , Figure 4 The images shown are images of surface damage to bridge steel structures and their segmentation and recognition results. The regions and distributions of surface damage to bridge steel structures can be effectively obtained from the images.

[0052] In particular, for images of rust, corrosion, and peeling, the original RGB color space image can be converted into HSV color space. The HSV parameters are adjusted to determine the color range between the diseased area and other areas. The color range between the diseased area and other areas is converted into a black and white mask. Identification is achieved through region recognition and extraction, thereby eliminating the manual labeling process for such diseases and improving the efficiency of identification and detection.

[0053] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A bridge steel structure disease visual image data processing method, characterized in that, Includes the following steps: Step 1: Enhancement and optimization of bridge steel structure defect image data. Generative adversarial network is used to learn from the existing bridge steel structure defect visual image dataset and expand it to generate a bridge steel structure defect dataset. Step 2: Multi-label image classification of bridge steel structure defects by integrating prior knowledge Based on prior knowledge, image classification and labeling of bridge steel structure defects are completed. A label feature learning network composed of two-layer graph convolutional networks is used to learn the latent semantic associations and co-occurrence rules between different defect labels, and output the semantic feature vector of label relationship. At the same time, a multi-axis self-attention network based on MaxViT is used to capture global contextual information in the image and output the high-dimensional feature vector of the image. The visual feature vector extracted from the image and the semantic feature vector learned from the label relationship are fed into a multimodal feature fusion module for deep fusion, and then the classification result is output by a classifier, and the result corresponds to the probability of the defect category of the image. Step 3: Image segmentation of bridge steel structure defects Construct a multi-scale segmentation network consisting of symmetrically configured encoding networks, decoding networks, connection networks located in the encoding and decoding networks, and a synthesis network; The encoder network obtains multi-scale disease features by compressing layer by layer through convolutional and pooling layers to capture contextual information, while the decoder network restores the feature scale layer by layer through upsampling and convolutional layers to accurately locate the position of each pixel. The synthesis network extracts the fusion feature output of the decoding unit and uses the spatial attention module to extract the spatial fusion features. The two output features are mapped to a probability map, and finally the OTSU algorithm is used to complete the binarization segmentation.

2. The method for processing visual image data of bridge steel structure defects according to claim 1, characterized in that, Step one, enhancing and optimizing image data of bridge steel structure defects, specifically includes: a. Collect images of bridge steel structure defects and perform necessary preprocessing, including grayscale unification and filtering for noise reduction; establish an original dataset of bridge steel structure defect images. ; b. Augmenting bridge steel structure defect images from the original dataset of beam steel structure defect images using DCGAN; the DCGAN adds a convolutional neural network to the adversarial generative network based on the traditional generator and discriminator; the loss function of the generative network is designed as follows: Where n is the total amount of data in the dataset; This refers to pseudo-images generated by a generative network; The probability that the current image is identified as a fake image. For random noise; design the loss function of the discriminant network as follows: ; c. Merge the generated images with the original images to construct a dataset of bridge steel structure defects.

3. The bridge steel structure disease visual image data processing method according to claim 2, characterized in that, In step a, when the original data comes from different acquisition terminals and configuration parameters, adaptive cropping and scaling of all images are completed according to the preset image aspect ratio, while unifying the range of image pixel values ​​to ensure the consistency of the original data.

4. The bridge steel structure disease visual image data processing method according to claim 1, characterized in that, In step three, image segmentation of bridge steel structure defects... The coding network consists of m+1 ascending connected coding units. The decoding network consists of m decoding units that are symmetrically arranged with respect to the first m encoding units and connected in reverse order. The network connection includes setting the kth ( m connection units between the m-th encoding unit and the (m-k+1)th decoding unit ; The encoding unit consists of a 3x3 convolutional layer and a 2x2 pooling layer with a stride of 2, where the i-th ( ) coding units The input is the (i-1)th encoding module. The output feature map is obtained by downsampling by 0.5 times; the decoding unit packet consists of a 2x2 upsampling layer with a stride of 2 and two 3x3 convolutional layers; the connection unit consists of two 3x3 convolutional layers. The first decoding unit is obtained by concatenating the outputs of the m-th and (m+1)-th encoding units and then extracting and fusing them through the connection unit. The kth ( The input of the (k-1)th decoding unit is the output of the (k-1)th decoding unit and the output of the (k)th encoding unit, which are then concatenated and extracted and fused by the connection unit.

5. The method for processing visual image data of bridge steel structure defects according to claim 4, characterized in that, The synthesis network obtains the fused feature output of the m-th decoding unit. and output the fused features The spatial fusion feature output is obtained by using the spatial attention module as input. The two fused feature outputs are finally processed by convolution compression and normalization and mapped to a channel with the same number of disease types to generate a defect prediction probability mapping map. Finally, the OTSU algorithm is used to complete the binarization segmentation.