A Deep Learning Super-Resolution Method for Crack Detection

By building a lightweight deep learning super-resolution network, the problems of difficulty in low-resolution image mapping and high computing resource utilization in crack detection are solved, and high-precision super-resolution amplification of crack images is achieved, texture information is retained, and visual experience is improved.

CN114612306BActive Publication Date: 2025-07-29BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210250155.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2025-07-29
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

The existing deep learning super-resolution technology has poor effect in crack detection, it is difficult to map low-resolution images to high-resolution images, and the computing resources are high. The existing methods lose texture information, affecting the visual experience.

Method used

A lightweight deep learning super-resolution network based on crack image features is constructed, and a lightweight residual module with a post-up sampling structure and attention mechanism is adopted. The head, Body, and Tail modules are built to perform end-to-end learning, reducing computing resource occupation and retaining texture information.

Benefits of technology

High-precision super-resolution amplification of crack images under low computing resources, retaining the texture information of cracks, improving the visual experience, and suitable for image preprocessing for crack detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114612306B_ABST
    Figure CN114612306B_ABST
Patent Text Reader

Abstract

The present invention discloses a deep learning super-resolution method for crack detection, belonging to the technical field of image super-resolution. The present invention includes the following steps: constructing a crack image dataset for super-resolution network training; constructing a super-resolution network for cracks; training the super-resolution network for cracks; super-resolution magnification of crack images. The present invention makes full use of the advantages demonstrated by deep learning in the field of image super-resolution. Based on the crack image features, a lightweight residual module including an attention mechanism and depthwise separable convolution is designed, and a super-resolution network is constructed using a post-upampling structure, solving the problems of difficult and inaccurate mapping of low-resolution crack images to high-resolution images, and performing super-resolution magnification on crack images with low computational resource occupancy, retaining the texture information of cracks and enhancing the visual experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image super-resolution, and particularly to a deep learning super-resolution method for crack detection. Background Art

[0002] At present, in the field of crack detection, it is mainly divided into manual subjective detection, using acoustic emission instruments and laser scanners for detection. The former method is restricted by the subjective consciousness of the detector and does not have universality. The latter method uses instrument equipment with high costs, which is not conducive to large-scale deployment. The rapid development of deep learning image processing technology has brought new opportunities for crack detection. Using this method to perform intelligent analysis on crack images can improve the efficiency of detection work and the accuracy of detection results, and can also reduce the workload of detection staff as much as possible and reduce detection costs. Processing images using deep learning is mostly completed based on convolutional neural networks. This network extracts features by continuously iterating the convolutional layer and then obtains the desired result through the mapping layer. Through most experiments, it is proved that this method indeed has advantages that cannot be compared with traditional methods and has good recognition performance for concrete cracks, road cracks, and rock layer cracks.

[0003] The commonly used detection network for crack detection is a semantic segmentation network. This network discriminates and classifies each pixel point of the image one by one. The more pixel points there are, the richer the content that the network can learn. Therefore, this has relatively high requirements for the resolution of the input image. In the authoritative datasets Concrete Crack Images for Classification and Crack-detection in the field of crack detection, the image resolution is about 224×224, while the image resolution required for input into the segmentation network is 480×480, 640×640 or higher. Therefore, it is necessary to magnify low-resolution images that do not meet the input requirements to obtain high-resolution images that meet the network requirements (low resolution and high resolution are relative. Compared with 200×200, 400×400 belongs to a high-resolution image, and compared with 600×600, it belongs to a low-resolution image). The commonly used magnification method is to interpolate and magnify the low-resolution image through bicubic interpolation. This method is simple and fast and does not require adding additional modules. This method is widely used for magnifying images on mobile phones and computers. However, the "coarse high-resolution" image obtained by rough interpolation magnification only has an improvement in resolution and will lose the texture information of the objects in the image. The most common problem is that the edges of the magnified objects are blurred, which affects the visual experience and is also not conducive to subsequent recognition and segmentation processing.

[0004] In view of the development of deep learning super-resolution technology, using a super-resolution network to achieve image magnification has become a popular research direction. This method learns from a large amount of data to fit a mapping relationship from low-resolution images to high-resolution images, which can not only improve the resolution of the image but also bring a better visual experience. It is an important image processing technology. However, the existing deep learning super-resolution technology mainly faces actual scene task types, such as landscapes, animals, plants, people, food, buildings, and vehicles. The mapping relationships constructed by existing networks work well in these scenarios, but when applied to crack-type images, distortion and blurring will occur, indicating that the previously constructed mapping relationships are not suitable for crack images. In addition, the existing super-resolution technology occupies a relatively high amount of computing resources and is not suitable as a preprocessing module for crack detection. Considering the above problems, the present invention designs a deep learning super-resolution method for crack detection, effectively solving the problems of difficult and inaccurate mapping from low-resolution crack images to high-resolution images, and performing super-resolution magnification on crack images while occupying less resources, retaining the texture information of the cracks and enhancing the visual experience. Summary of the Invention

[0005] The main technical problems to be solved by the present invention are that the existing deep learning super-resolution technology has poor effects, it is difficult to map low-resolution crack images to high-resolution images, and the computing resources occupied are relatively high. Therefore, a deep learning super-resolution method for crack detection is constructed to improve the resolution of crack images while retaining the texture information of the cracks and reducing the computing resources. To achieve the above object, the present invention adopts the following technical solutions:

[0006] A deep learning super-resolution method for crack detection, comprising the following steps:

[0007] Step 1: Construct a crack image dataset for network training.

[0008] In deep learning, the quality of the dataset is crucial for the result of super-resolution. Therefore, an original crack image dataset is constructed through publicly available network image data and field-collected image data, and then through data augmentation, data clipping, and data downsampling, a crack image set for training and supervision is constructed. Such operations not only maintain the high reliability of the dataset but also have a large scale.

[0009] Step 2: Construct a super-resolution network for cracks.

[0010] Based on the post-upsampling super-resolution network structure, a lightweight super-resolution network for cracks is constructed. This structure performs end-to-end learning on low-resolution images and adds a learnable upsampling layer at the end to fit high-resolution images. The advantage of this is that it can greatly reduce the occupation of computing resources. The designed network can be divided into three modules: Head, Body, and Tail.

[0011] The Head module consists of two ordinary convolutional layers, which are used to increase the dimension of the input low-resolution image and extract preliminary texture information.

[0012] Furthermore, the Body module consists of 16 repeatedly stacked Blocks and 1 convolutional layer, which are used for refined texture feature extraction. Each Block is divided into three parts: the front, middle, and rear. The front part uses an ordinary convolutional layer to extract the information output by the previous layer. The middle part is a lightweight residual structure containing an attention mechanism and depthwise separable convolutions, and feature fusion is performed between the input and output using a skip connection. Finally, another ordinary convolutional layer is used to summarize the texture information in the middle part. Finally, the input and output of the Body module are fused again using a skip connection to enhance the interaction between texture information.

[0013] The final Tail module consists of two ordinary convolutional layers and one sub-pixel convolutional layer, which are used to perform upsampling on the feature map output by the Body to achieve the magnification of the low-resolution image.

[0014] Step 3: Train the super-resolution network for cracks.

[0015] Input the constructed crack super-resolution dataset into the designed network, select the L1 loss function and the Adam optimizer, and train the network to fit the mapping relationship between the low-resolution crack image and the high-resolution crack image. After training, save the model with the highest fitting rate.

[0016] Step 4: Super-resolution magnification of crack images.

[0017] Use the trained model file to map the low-resolution crack image to obtain the magnified high-resolution crack image. Save the image, which can be used for subsequent crack detection.

[0018] The present invention has the following advantages compared with the prior art:

[0019] 1. The present invention is a deep learning super-resolution method designed based on the characteristics of crack images, effectively solving the problems of distortion and blurring of the magnified crack images in the existing methods, having a good visual experience and crack texture information, and can be used for image preprocessing in crack detection and recognition tasks.

[0020] 2. A lightweight residual module containing an attention mechanism and depthwise separable convolutions is designed and a lightweight super-resolution network is constructed using a post-upsampling structure, effectively reducing the occupation of computing resources while ensuring that the network can perform high-precision super-resolution. Description of the Drawings

[0021] Figure 1 It is a schematic diagram of the overall process of the deep learning super-resolution method for crack detection in the present invention.

[0022] Figure 2 It is a schematic diagram of the process for constructing a crack image dataset for training in the present invention.

[0023] Figure 3 It is a structural diagram of the super-resolution network for crack detection in the present invention. Specific embodiments

[0024] The present invention mainly realizes the super-resolution of crack images. The specific method adopted by the present invention will be introduced in detail below with reference to the accompanying drawings.

[0025] Specifically, the process of the deep learning super-resolution method for crack detection is as shown in Appendix Figure 1 and includes the following steps. S1: Construct a crack image dataset for network training. S2: Construct a super-resolution network for cracks. S3: Train the super-resolution network for cracks. S4: Super-resolution magnification of crack images.

[0026] (1) For S1: Construct a crack image dataset for network training.

[0027] The process is shown in Appendix Figure 2 . Obtain some crack images through public resources, and use cameras to take cracks of different categories in different scenarios as supplements to improve the diversity of crack categories. Randomly rotate and supplement light to the data to improve the generalization of the data. Then crop the processed image data to obtain supervised images as the ground truth; then process the cropped images by downsampling and adding noise to generate low-resolution images for training.

[0028] (2) For S2: Construct a super-resolution network for cracks.

[0029] The overall structural diagram of the network is shown on the left side of Appendix Figure 3 . The designed network can be divided into three modules: Head, Body, and Tail, where relu is the activation function.

[0030] The Head module consists of two ordinary convolutional layers with a convolutional kernel size of 3, expands the number of channels of an image with an input resolution of 96×96 and 3 channels to 64, and performs preliminary feature extraction.

[0031] The Body module consists of 16 repeatedly stacked Blocks and 1 ordinary convolutional layer, and a skip connection is made between the input and output of the Body for fine-grained texture feature extraction. The specific structure of the Block is shown in Appendix Figure 3On the right side. Each Block is divided into three parts: front, middle, and back. The front part uses a normal convolution with a kernel size of 3 to extract information from the output of the previous layer.

[0032] The middle part is a lightweight residual structure, that is, there is a skip connection between the input and the output for feature information fusion. This residual structure first uses a normal convolution with a kernel size of 1 to increase the number of image channels to 128, and then uses a depthwise separable convolution with a kernel size of 3 for information extraction. This convolution sets the number of kernel channels to 1 and the number of kernels to the number of feature map channels, so that each kernel channel can process each feature map channel. Compared with the use of multi-channel convolution kernels to process feature maps in normal convolution, depthwise separable convolution highly optimizes matrix multiplication and has very low parameter quantities. After passing through the depthwise separable convolution, a channel separation operation is performed on the feature map, and the feature map is divided into two feature maps with 64 channels and the same size, which are respectively sent into two branches: channel attention (left) and spatial attention (right). Among them, the channel attention performs average pooling on each channel of the feature map to obtain a one-dimensional vector; then, through two normal convolution layers with a kernel size of 1, an output vector is obtained. This vector analyzes the weight relationship for each channel of the feature map and assigns a greater weight to the more important channels; finally, the output vector is normalized and then multiplied by the input feature map to obtain a new feature map. The spatial attention focuses on the spatial information of the feature map, performs a mean operation on each pixel in all channels of the feature map to obtain a single-channel feature map; then, through a normal convolution layer with a kernel size of 7, an output feature map is obtained; finally, the output feature map is normalized and then multiplied by the input feature map to obtain a new feature map. These two attention mechanisms can effectively assign corresponding weights to the feature map channels and pixels according to the region of interest of the feature map, improving the feature extraction ability. Then, the feature maps processed by these two attention mechanisms are subjected to a channel concatenation operation, and then a normal convolution with a kernel size of 1 is used to reduce the number of channels and restore the state of the input feature map of this residual structure.

[0033] The back part is a normal convolution with a kernel size of 3, which summarizes the information output by the middle residual structure. Finally, the input and output of the Block module are fused in a skip connection manner to enhance the interaction between texture information.

[0034] The Tail module consists of two ordinary convolutional layers with a kernel size of 3 and a sub-pixel convolutional layer, which are used to perform upsampling operations on the feature maps output by the Body to achieve the magnification of low-resolution images. As a commonly used image super-resolution upsampling module, sub-pixel convolution can well increase the size of the image while retaining the original information of the image. The network designed with this structure takes into account the two advantages of low computational complexity and high accuracy, and well completes the task of super-resolution magnification of crack images.

[0035] (3) For S3: Train the super-resolution network for cracks.

[0036] Input the constructed crack super-resolution dataset into the designed network, select the L1 loss function and the Adam optimizer, set the number of training times to 200, the initial learning rate to 0.0005, the size of the input low-resolution training image to 96x96, the size of the supervised image to 192×192, calculate the loss between the image after the training image enters the network and the supervised image, use the adaptive learning rate to find a balance between the training speed and accuracy, and continuously optimize the network through the optimizer to fit the mapping relationship between the low-resolution image and the high-resolution image. After the training is completed, save the model with the highest fitting rate.

[0037] (4) For S4: Super-resolution magnification of crack images.

[0038] Use the trained model file to map the low-resolution crack images to obtain high-resolution crack images. The high-resolution images obtained by this network have better texture detail information and visual experience compared with the images magnified by the traditional interpolation method. Finally, save the images, which can be used for subsequent crack detection.

[0039] The above specific implementation manners are only used to illustrate the technical solutions of the present invention, rather than to limit it. Those skilled in the art should understand that: the above implementation manners do not limit the present invention in any form, and all similar technical solutions obtained by using equivalent replacements or equivalent transformations and the like belong to the protection scope of the present invention.

Claims

1. A deep learning super-resolution method for crack detection, characterized in that It includes the following steps: S1: Construct a crack image dataset for network training; S2: Construct a super-resolution network for cracks; S3: Train the super-resolution network for cracks; For S2, construct a super-resolution network for cracks; the designed network can be divided into three modules: Head, Body, and Tail, where relu is the activation function; the Head module consists of two ordinary convolutional layers with a convolutional kernel size of 3, which expands the number of channels of an image with an input resolution of 96×96 and 3 channels to 64 and performs preliminary feature extraction; the Body module consists of 16 repeatedly stacked Blocks and 1 ordinary convolutional layer, and a skip connection is made between the input and output of the Body for fine-grained texture feature extraction; each Block is divided into three parts: the front, middle, and back. The front part uses an ordinary convolutional layer with a convolutional kernel size of 3 to extract the information output by the previous layer; the middle part is a lightweight residual structure, that is, a skip connection is made between the input and output for feature information fusion; this residual structure first increases the number of image channels to 128 through an ordinary convolutional layer with a convolutional kernel size of 1, and then performs information extraction through a depthwise separable convolutional layer with a convolutional kernel size of 3. The depthwise separable convolutional layer sets the number of convolutional kernel channels to 1 and the number of convolutional kernels to the number of feature map channels. After depthwise separable convolution, a channel separation operation is performed on the feature map, and the feature map is divided into two feature maps with 64 channels and the same size, which are respectively sent into two branches of channel attention and spatial attention; Among them, channel attention performs average pooling on the feature map of each channel to obtain a one-dimensional vector; then, through 2 ordinary convolutional layers with a convolutional kernel size of 1, an output vector is obtained. This output vector analyzes the weight relationship of each channel of the feature map and assigns a greater weight to more important channels; finally, the output vector is normalized and then multiplied by the input feature map to obtain a new feature map; Spatial attention focuses on the spatial information of the feature map, performs average pooling on all channels of the feature map to obtain a single-channel feature map; then, through an ordinary convolutional layer with a convolutional kernel size of 7, an output feature map is obtained; finally, the output feature map is normalized and then multiplied by the input feature map to obtain a new feature map; then, the feature maps processed by these two attention mechanisms are subjected to a channel concatenation operation, and then through an ordinary convolutional layer with a convolutional kernel size of 1, the number of channels is reduced to restore the state of the input feature map of this residual structure; the back part is an ordinary convolutional layer with a convolutional kernel size of 3, which summarizes the information output by the middle residual structure. Finally, the input and output of the Block module are fused in a skip connection manner to enhance the interaction between texture information; The Tail module consists of two ordinary convolutional layers with a convolutional kernel size of 3 and a sub-pixel convolutional layer for upsampling the feature map output by the Body; For S3, train a crack-oriented super-resolution network; input the constructed crack super-resolution dataset into the designed network, select the L1 loss function and the Adam optimizer, set the number of training times to 200, the initial learning rate to 0.0005, the size of the input low-resolution training image to 96x96, and the size of the supervised image to 192×192. Calculate the loss between the image after the training image enters the network and the supervised image and optimize the fitting relationship. After training, save the model with the highest fitting rate; use the trained model file to map the low-resolution crack image to obtain the high-resolution crack image.

2. A deep learning super-resolution method for crack detection according to claim 1, wherein: For S1, construct a crack image dataset for network training; obtain some crack images through public resources, and use a camera to take cracks of different categories in different scenarios as a supplement to improve the diversity of crack categories; randomly rotate and supplement light to the data to improve the generalization of the data; then crop the processed image data to obtain a supervised image as the true value; then process the cropped image by downsampling and adding noise to generate a low-resolution image for training.

Citation Information

Patent Citations

  • Image crack segmentation method based on full convolutional neural network

    CN111028217A

  • Image super-resolution reconstruction method based on fused attention mechanism residual network

    CN111192200A