A real-time detection method and device for bridge concrete cracks based on edge computing and Transformer
By using a lightweight network based on Transformer on edge devices to detect bridge concrete cracks, the problem of excessive computing resources in the prior art is solved, and efficient edge computing detection is achieved.
Patent Information
- Application Number
- CN202210577573.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-25
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-05-25
AI Technical Summary
The existing bridge concrete crack detection methods require a large amount of computing resources and cannot be operated on edge equipment, resulting in inefficient detection.
Real-time detection of bridge concrete cracks is used based on edge computing and Transformer, and the amount of model parameters is reduced through the hybrid block transformation method to make it run on edge devices.
The detection of bridge concrete cracks that operate efficiently on edge equipment is realized, which improves detection efficiency and reduces the calculation load.
Smart Images

Figure CN114937016B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of crack detection and edge computing, and specifically relates to a real-time detection method and device for bridge concrete cracks based on edge computing and Transformer. Background Art
[0002] The surface of the concrete structure of bridges often has cracks of varying degrees. In addition to extreme natural disasters, uneven stress caused by long-term overload, deformation caused by temperature changes, and structural corrosion caused by long-term erosion by air and rain are also important factors in the generation of cracks. Cracks can accelerate the deterioration process of the surface structure. If these cracks are discovered in the early stages, the damage can be further reduced, avoiding casualties and economic losses.
[0003] The early detection and maintenance of cracks in bridge concrete structures mainly rely on manual inspection, which requires a lot of manpower and material resources. In addition, inspectors cannot detect all areas of the bridge, resulting in insufficient crack detection and posing a hidden danger to bridge safety. In recent years, the development of drone technology has greatly improved the detection efficiency of bridge concrete cracks. Drone technology uses airborne cameras to capture structural surface images of bridges and uses crack detection algorithms in cloud services to detect cracks in the captured images. Drone-based bridge inspection has attracted the attention of many researchers and bridge maintenance departments due to its safety and reliability. However, with the increase in the number of drones, the computing pressure of cloud services has also increased, limiting the efficiency of drone bridge inspection.
[0004] Edge computing is a decentralized computing architecture that breaks down large services that were originally handled entirely by central nodes into smaller, more manageable parts and distributes them to edge nodes for processing. This technology has become a new option for drone bridge inspection. Recently, Transformer has achieved excellent results in the field of image recognition and has gradually become the new mainstream of visual modeling. However, the current Transformer parameter volume and computing requirements are too large to run on edge devices. The deep learning algorithm in drone edge computing must have fewer parameters to effectively detect cracks. Summary of the invention
[0005] The purpose of the present invention is to solve the problem that the existing bridge concrete crack detection method requires too much computing resources and cannot be run on edge devices, and to provide a real-time detection method and device for bridge concrete cracks based on edge computing and Transformer.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] On the one hand, a real-time detection method for bridge concrete cracks based on edge computing and Transformer is provided, including:
[0008] Step S1: Collecting bridge concrete crack images; including collecting public crack images and using drones to take crack images;
[0009] Step S2: Check and preprocess the images to establish a data set; check the images taken by the drone, and use the fuzzy denoising method to process the blurry images; use cropping, color dithering, and scaling to expand the data; manually annotate the images taken by the drone and divide them into two categories: images with cracks and images without cracks, and merge the crack images publicly available on the Internet with the crack images taken by the drone to construct a data set;
[0010] Step S3: construct a lightweight Transformer network suitable for edge computing; the network includes three parts: input conversion, Transformer encoder and prediction head, wherein the input conversion includes block division, hybrid block transformation and linear projection, the image needs to be divided into 8*8 image blocks before inputting the network, and then the hybrid block transformation is performed, and the input sequence of the Transformer encoder is formed after linear projection, and the embedding dimension of each block is 128; position coding needs to be inserted between the input conversion and the Transformer encoder; the Transformer encoder mainly includes a multi-head attention layer (2 heads) and a feedforward network (FFN), wherein the first layer dimension of the FFN is 512, the second layer dimension is 128, and the activation function is GELU; the prediction head includes a global average pooling layer and a linear layer, wherein the linear layer is a two-layer perceptron, the input dimension is 128, the first layer output dimension is 512, the second layer is used for classification, and the output dimension is the number of categories 2, representing crack-free images and cracked images respectively;
[0011] Step S4: training the network to obtain a recognition model; pre-training the lightweight Transformer network on the large-scale image dataset ImageNet-1K, and then fine-tuning the pre-trained model on the bridge concrete crack dataset obtained in step S2;
[0012] Step S5: Deploy the recognition model to the drone and use the drone to perform real-time detection of bridge concrete cracks; deploy the trained lightweight Transformer model to the AI development board; formulate a calculation offloading plan based on the model's computing power, transmission latency and other information; mount the AI development board on the drone as an edge computing device, and when the drone performs real-time detection, it will transfer the offloaded bridge concrete crack detection task to the edge computing device for calculation, and transmit the identified category and confidence back to the ground device.
[0013] Optionally, in a possible implementation, the hybrid block transform method includes two types of transform methods, as follows:
[0014] The first type of transformation is block discarding transformation, in which half of the non-adjacent blocks are discarded from all the blocks obtained after segmentation. Let P be the block matrix after segmentation. The main formula of the block discarding method is as follows:
[0015] P=P*D
[0016] Where P has i rows and j columns, and D is a binary matrix (0 or 1), where 0 represents the block to be discarded and 1 represents the block to be retained; the binary matrix D is calculated by the following formula:
[0017] D = CycleAndReshape(0,1,n)
[0018] First, a binary array is generated, which is composed of alternating cycles of 0 and 1. n = i*j represents the total number of image blocks and is the length of the generated binary array. Then, the binary array is reconstructed into a matrix of the same size as P. Finally, the P matrix is multiplied by the D matrix, and the resulting matrix is the transformed block matrix.
[0019] The second type of transformation is block thumbnail transformation, which randomly selects half of all image blocks and converts them into thumbnails, with the size converted from the original 8*8 to 4*4; these thumbnails are mixed with the other half of the image blocks in pairs, and only the mixed image blocks are retained. The mixing formula is as follows:
[0020] patch mix =Mpatch1+Fill(Tn(patch2))
[0021] Among them, patch mix is a new mixed sample block, M is a binary mask representing the area where the image block is cropped and retained, the Fill function fills and generates an image block of the same size as patch2, and the Tn function is a thumbnail operation;
[0022] The hybrid block transform is a method used in the model training stage, which transforms the input image blocks during training, discards part of the input information, and allows the Transformer to use an incomplete input sequence for training, and uses the complete input sequence for calculation during inference; during training, the hybrid block transform uses a transformation method to transform the input image blocks each time, and the transformation method used can be controlled by a manually set probability threshold according to actual needs. On the other hand, the present invention also provides a real-time detection device for bridge concrete cracks based on edge computing and Transformer, including:
[0023] An acquisition module, which acquires images of bridge concrete cracks;
[0024] The data set module performs data inspection and preprocessing on the images and establishes the data set;
[0025] Network module, building a lightweight Transformer network suitable for edge computing;
[0026] Training module, training the network to obtain the recognition model;
[0027] The detection module deploys the recognition model onto the drone and uses the drone to conduct real-time detection of bridge concrete cracks.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] In the hybrid block transformation method used in model training, the first type of transformation alleviates the problem of Transformer overfitting for small data sets and improves the operating efficiency of Transformer. In the second type of transformation, half of the randomly selected image blocks are converted into thumbnails and mixed with the other half of the image blocks in pairs, which enhances the diversity between blocks. The thumbnails of image blocks with cracks can enhance the model's learning of subtle cracks, and the thumbnails of image blocks without cracks can be used as occlusion to improve the generalization ability of the model. In addition, the lightweight Transformer reduces the computational load, allowing it to be deployed on resource-constrained platforms such as embedded devices and run efficiently in edge computing devices supported by drones. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a flow chart of the real-time detection method of bridge concrete cracks based on edge computing and Transformer;
[0031] Figure 2 This is the architecture diagram of the lightweight Transformer;
[0032] Figure 3 An architectural diagram for real-time detection of drones equipped with edge computing devices;
[0033] Figure 4 This is a flow chart of the real-time detection device for bridge concrete cracks based on edge computing and Transformer. DETAILED DESCRIPTION
[0034] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application are described below. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present application.
[0035] It should be understood that in various embodiments of the present invention, the size of the sequence number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0036] It should be understood that in the present invention, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.
[0037] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0038] The present invention provides a real-time detection method for bridge concrete cracks based on edge computing and Transformer. Figure 1 As shown in the flowchart, the specific process includes:
[0039] Step S1: collecting bridge concrete crack images;
[0040] S101, collecting public crack images;
[0041] A bridge concrete crack dataset published in the literature “Automated Bridge Crack Detection Using Convolutional Neural Networks” was collected. The dataset includes 4856 training images and 1213 test images, both with a resolution of 224*224, and a ratio of crack images to non-crack images of approximately 2:1.
[0042] S102, using a drone to take images of the cracks;
[0043] First, develop a flight plan for the drone. According to the differences in the inspection components, formulate different safety distances: bridge piers and towers are generally controlled at about 1.5-3m; if the bridge has complex components such as cables, it should be controlled at about 3-5m. The specific safety distance needs to be determined in combination with the on-site conditions. During a single flight, only the upper and lower parts of one side of the bridge are inspected. The inspection order is from left to right, first up and then down. When inspecting the lower part, the upper camera is used to operate on the same side as the drone pilot. Before the drone flies, it is necessary to check the weather conditions and the drone function. Weather is an important factor affecting drone aerial photography. For example, when the wind speed exceeds the maximum range that the drone's flight control system can withstand, its stability will be difficult to maintain, the quality of aerial photography will be seriously affected, and in severe cases, the drone will be damaged. For shooting the surface of the bridge, it should be carried out when there is sufficient light and high air visibility to avoid backlighting. In case of bad weather, it is necessary to wait until the weather conditions are more favorable before flying.
[0044] Before takeoff, the drone needs to be placed steadily at the takeoff point determined by the flight plan. Debug each system to ensure normal operation at startup, confirm that the takeoff area is safe, and then unify the command to dispatch the takeoff. After takeoff, the drone flies on the left and right sides of the bridge cable respectively, and does not enter the area above the main body of the bridge to ensure that it does not affect the normal passage of vehicles on the bridge. During takeoff, the drone pilot controls the aircraft in real time through the remote control, and the observer observes the aircraft status through the parameters transmitted back by the aircraft. After the aircraft reaches a safe height, the pilot retracts the landing gear through the remote control, switches the flight mode to the automatic mission flight mode, and flies according to the route set in the flight plan. At the same time, the pilot needs to pay attention to the dynamics of the drone by visually observing the drone, switch to manual mode at any time to adjust the aircraft status according to the specific flight situation, and the observer pays attention to the battery status, flight speed, flight altitude, flight attitude, and route completion status displayed in the flight control software. After completing the crack image shooting task, land the drone at the predetermined location.
[0045] Step S2: Check and preprocess the image data to establish a data set;
[0046] S201, image inspection and preprocessing;
[0047] Import the images taken by the drone into the local device and check whether the images taken by the drone are blurry. Use the blur denoising method to process the blurry images. If the image quality after processing is still low, it can be excluded. Then, use cropping, color dithering, and scaling to expand the data.
[0048] S202, data set construction;
[0049] The images taken by the drone were manually annotated and divided into two categories: images with cracks and images without cracks. The crack images publicly available on the Internet were merged with the crack images taken by the drone, and all were converted into RGB images in PNG format and placed in two folders, one with cracks and one without cracks, to complete the dataset construction.
[0050] Step S3: Build a lightweight Transformer network suitable for edge computing;
[0051] The architecture of the lightweight neural network is as follows Figure 2 As shown, it includes three parts: input conversion, Transformer encoder and prediction head, wherein the input conversion includes patching, mixed block transformation and linear projection (LinearProjection&Reshape), the Transformer encoder mainly includes two parts: self-attention layer and forward network, and the prediction head includes global average pooling layer and linear layer; positional encoding needs to be inserted between the input conversion and the Transformer encoder;
[0052] The block division is to divide the input image into 8*8 image blocks; the hybrid block transformation is a method used in the model training stage, which transforms the input image blocks during training, discards part of the input information, and allows the Transformer to use an incomplete input sequence for training, and use the complete input sequence for calculation during inference. The hybrid block transformation includes two types of transformations. The first type of transformation is the block discarding transformation, which discards half of the non-adjacent blocks in all the blocks obtained after segmentation. Let P be the block matrix after segmentation, and the main formula of the block discarding method is as follows:
[0053] P=P*D
[0054] Where P has i rows and j columns, and D is a binary matrix (0 or 1), where 0 represents the block to be discarded and 1 represents the block to be retained. The binary matrix D is calculated by the following formula:
[0055] D = CycleAndReshape(0,1,n)
[0056] First, a binary array is generated, which is composed of alternating cycles of 0 and 1. n = i*j represents the total number of image blocks and is the length of the generated binary array. Then, the binary array is reconstructed into a matrix of the same size as P. Finally, the P matrix is multiplied by the D matrix, and the resulting matrix is the transformed block matrix.
[0057] The second type of transformation is block thumbnail transformation, which randomly selects half of all image blocks and converts them into thumbnails, with the size converted from the original 8*8 to 4*4. These thumbnails are mixed with the other half of the image blocks in pairs, and only the mixed image blocks are retained. Suppose there are two sample blocks patch1 and patch2, using thumbnail To replace a random area of patch1, but do not change the label, h and w represent the width and height of the thumbnail Tn(patch2), respectively, both of which are 4 in this embodiment. The transformation formula is as follows:
[0058] patch mix =Mpatch1+Fill(Tn(patch2))
[0059] Among them, patch mix is a new mixed sample block, M is a binary mask, representing the area where the image block is cropped and retained, and the Fill function fills and generates an image block of the same size as patch2; to obtain the binary mask M, the bounding box coordinates B of the cropped area on patch1 are required = (c x ,c y ,c w ,c h ) is sampled, then area B in patch1 is removed and filled with the thumbnail Tn(patch2). The coordinates of the bounding box are uniformly sampled, as shown in the following formula:
[0060] c x ~Unif(0,W),c y ~Unif(0,H)
[0061] c w =w,c h =h
[0062] Where W and H are the width and height of the original image block, respectively. This hybrid strategy enables the network to learn the same image at different scales. During training, the hybrid block transform uses a transformation method to transform the input image block each time. The transformation method used can be controlled by a manually set probability value according to actual needs.
[0063] After the hybrid block transformation, the input sequence of the Transformer encoder is formed by linear projection, and the embedding dimension of each block is 128. The Transformer adds an additional learnable [class] tag to the sequence of embedded image blocks, representing the classification parameters of the entire image, which contains potential information for classification.
[0064] The Transformer encoder includes a multi-head attention layer (Multi-head Attention Layers) (2 heads) and a position-based feedforward neural network (FFN), the first layer dimension of the FFN is 512, the second layer dimension is 128, and the activation function is GELU; the outputs of the multi-head attention layer and the feedforward network are sent to the "add and norm" layer for processing, which contains a residual structure and layer normalization.
[0065] The self-attention mechanism involved in the Transformer encoder is an important component of the Transformer. Compared with the traditional recurrent neural network (RNN) model, it can capture the "long-term" dependencies between sequence elements. The main difference between self-attention and convolution operations is that the weights are calculated dynamically, rather than static weights like convolution (the same for any input). In addition, self-attention can keep the arrangement and number of input points unchanged. Therefore, compared with standard convolution that requires a grid structure, it can easily operate on irregular inputs. In fact, self-attention provides the ability to learn global and local features, and provides the ability to adaptively learn kernel weights and receptive fields (similar to variable convolution). The self-attention layer updates each component of the sequence by aggregating the global information of the complete input sequence. Let Represents an input sequence of length n (x1, x1, x1…x n ), where d represents the embedding dimension of each input. The goal of the self-attention mechanism is to capture the interactions between all n inputs by encoding the global context information of each input. This is done by defining three learnable weight matrices, including the query vector Key Vector Sum value vector where d q =d k The input sequence X is multiplied by these weight matrices to obtain Q = XW Q , K=XW K and V=XW V The output of the self-attention layer It is obtained by the following formula:
[0066]
[0067] For an input in the sequence, the self-attention layer calculates the dot product of the query vector and the key vector, and then normalizes it with the Softmax operator to get the attention score. Then, the value of each input is multiplied by the attention score to get the weighted output.
[0068] In order to include various complex relationships between different elements in the sequence, the Transformer architecture further expands the single-head attention into a multi-head attention mechanism. The multi-head attention mechanism includes multiple self-attention blocks, each of which has a set of learnable weight matrices. For an input X, the outputs of the h self-attention blocks in the multi-head attention are concatenated into a single matrix And projected into the weight matrix The multi-head self-attention mechanism achieves the ability to focus on multiple specific locations by giving the attention layer different subspace representations, enriching the diversity of the feature subspace without additional computational cost.
[0069] Afterwards, the output of the multi-head attention layer is fed into a two-layer feed-forward neural network (FFN), which is represented as follows:
[0070] F2(GELU(F1(x)))
[0071] Where F1 and F2 are linear layers, and the function form is W x +b, the dimension of F1 is 512, and the dimension of F2 is 128.
[0072] The linear layer in the prediction head is a two-layer perceptron (MLP) with an input dimension of 128, an output dimension of 512 for the first layer, and a second layer for classification with an output dimension of 2, representing no cracks and cracks.
[0073] Step S4: training the network to obtain a recognition model;
[0074] The lightweight Transformer network is pre-trained on the large-scale image dataset ImageNet-1K, and then the pre-trained model is used to fine-tune the bridge concrete crack dataset obtained in step S2. The model parameters during fine-tuning are set as follows: batch_size is set to 32, the learning rate lr is set to 1e-4, cosine learning rate decay is used, the minimum learning rate min_lr is set to 1e-6, the Adam optimizer is used, the number of iterations is set to 100, and the probability of mixed block transformation is 0.5. The ratio of the training set to the test set is divided into 8:2. During model training, the training set images are randomly flipped left and right, randomly adjusted in brightness, randomly adjusted in contrast, and normalized. During testing, the test set images are only normalized.
[0075] The lightweight Transformer model is trained using the TensorFlow deep learning framework. The images of the training set are input into the lightweight Transformer network. In an iterative optimization process, 32 training samples are randomly sampled to form a batch. The network is optimized and the network parameters are updated in combination with the back propagation algorithm. After the training is completed, the model with the highest accuracy on the test set is saved to obtain the crack detection model (Pb file format).
[0076] Step S5: deploying the recognition model to the drone, and using the drone to perform real-time detection of bridge concrete cracks;
[0077] S501. Deploy the trained lightweight Transformer model to the AI development board.
[0078] The NVIDIA Jetson TX2 development board is used as the deployment platform for the Transformer model. The NVIDIA Jetson TX2 is a high-performance, low-power computing module with a core size of only 87mm×50mm, which is suitable for deploying deep learning technology on small cutting-edge devices such as drones. The NVIDIA Jetson TX2 uses a 256-core NVIDIA Pascal architecture GPU, a dual-core Denver 2 64-bit CPU + a quad-core ARM A57 MPCore, built-in 8GB 128-bit LPDDR4 memory, 32GB eMMC 5.1 storage, and supports 802.11ac WiFi + Bluetooth and 10 / 100 / 1000BASE-T adaptive Ethernet. The NVIDIA Jetson TX2 is pre-equipped with the JetPack embedded application development SDK, allowing developers to use the NVIDIA DIGITS interface to train deep learning models in the cloud, data center or computer.
[0079] First, you need to use Jetpack to install the operating system and necessary deep learning libraries for NVIDIA Jetson TX2, including TensorRT, cuDNN, CUDA Toolkit, etc. The operating system chooses to install Ubuntu 16.04. Then convert the lightweight Transformer Pb model trained by TensorFlow in step S4 to an uff model available for TensorRT. TensorRT is a high-performance deep learning inference optimizer developed by NVIDIA, which can provide low-latency and high-throughput deployment inference for deep learning applications. This framework can parse the TensorFlow model, then map it one-to-one with the corresponding layers in TensorRT and convert it to TensorRT. Then, you can implement optimization strategies for NVIDIA's GPU in TensorRT and accelerate deployment. Call the convert script that comes with the uff package to complete the model conversion.
[0080] Then use TensorRT to deploy the converted uff model. First, import the weights and network structure saved in the uff model, and then execute the optimization algorithm to generate the corresponding TensorRT inference engine. Before creating the engine, you need to specify the maximum batch size. After the inference engine is generated, you can input the image to be predicted for inference and get the predicted output. After testing the inference process and output results on the development version, the model deployment of the AI development board can be completed, and it will be used as an edge computing node for subsequent detection.
[0081] S502, formulating a plan for computing offloading;
[0082] Due to the need for real-time crack detection, an offloading decision is made with the goal of reducing latency, so that the total time delay is minimized while ensuring that the detection frame rate is no less than 24FPS. The time delay includes the transmission time of the offloaded data (the video frame of the bridge surface taken by the drone) to the edge computing node, the detection time at the computing node, and the transmission time of the ground-end device receiving the calculation results from the edge computing node (including the judgment result of whether it is a crack and the confidence level). The Markov decision process is used to analyze the average latency of the task execution and the average power consumption of the device, and whether to perform the offloading is determined based on the available bandwidth, the size of the data to be offloaded, and the energy consumption of the task. For the task of detecting cracks in bridge concrete, a binary-based offloading method is used to offload it as a whole to the edge computing device for execution.
[0083] After the offloading decision is made, resource allocation is required. The offloaded data is first transmitted to the scheduler in the edge computing device. The scheduler checks whether there are sufficient computing resources. If so, it calls the edge node to process the detection task. If the computing power provided by the edge computing node is insufficient, the crack detection task is delegated to the cloud for processing.
[0084] S503, using a drone equipped with an AI development board to perform real-time detection;
[0085] The NVIDIA Jetson TX2 development board is mounted on the drone as an edge computing device. Through the cooperation of the ground-side program, the drone flight control program, and the airborne edge computing program, the drone can inspect the set area and detect the cracks in the concrete of the bridge using AI. Figure 3 As shown. According to step S102, the drone real-time detection task is performed. During the flight of the drone, the camera system carried by it will obtain the video stream image data of the bridge surface, and transmit the data stream to be unloaded to the edge computing device through the calculation offloading scheme formulated in step S502. Then, the AI development board carried by the drone acts as an edge computing server, using the lightweight Transformer model deployed therein for detection, and transmits the recognition result of whether the video frame image is a crack and the confidence back to the ground station. Finally, the ground-end device displays the detected category and confidence in real time in the received video stream.
[0086] In addition, unlike ground servers, which do not need to consider energy consumption, the edge computing devices on drones provide computing services for ground stations, and the energy consumption of drones also needs to be considered. During the flight, drones need to adjust their flight plans according to actual energy consumption and return home in time when the battery is about to run out.
[0087] The present invention provides a real-time detection device for bridge concrete cracks based on edge computing and Transformer, such as Figure 4 Its structure diagram shown includes:
[0088] An acquisition module, which acquires images of bridge concrete cracks;
[0089] The data set module performs data inspection and preprocessing on the images and establishes the data set;
[0090] Network module, building a lightweight Transformer network suitable for edge computing;
[0091] Training module, training the network to obtain the recognition model;
[0092] The detection module deploys the recognition model onto the drone and uses the drone to conduct real-time detection of bridge concrete cracks.
[0093] Furthermore, the acquisition module includes two parts: collecting public crack images and using drones to photograph crack images. The use of drones to photograph crack images specifically includes:
[0094] Develop a flight plan for the drone, and formulate different safety distances according to the differences in the inspection components: bridge piers and towers are generally controlled at about 1.5-3m; if the bridge has complex components such as cables, it should be controlled at about 3-5m. The specific safety distance needs to be determined in combination with the on-site conditions. During a single flight, only the upper and lower parts of one side of the bridge are inspected. The inspection order is from left to right, first up and then down. When inspecting the lower part, the upper camera is used to operate on the same side as the drone pilot. The drone needs to be inspected before flight, including weather conditions and drone function checks. Before takeoff, the drone needs to be placed stably at the takeoff point determined by the flight plan. Debug each system to ensure normal operation at the start-up, confirm that the takeoff area is safe, and then unify the command to dispatch the takeoff. After takeoff, the drone flies on the left and right sides of the bridge cable respectively, and does not enter the area above the main body of the bridge to ensure that it does not affect the normal passage of vehicles on the bridge. During takeoff, the drone pilot controls the aircraft in real time through the remote control, and the observer observes the aircraft status through the parameters transmitted back by the aircraft. After the aircraft reaches a safe height, the pilot retracts the landing gear through the remote control, switches the flight mode to the automatic mission flight mode, and flies according to the route set in the flight plan. At the same time, the pilot needs to pay attention to the dynamics of the drone by visually observing the drone, and switch to manual mode at any time to adjust the aircraft status according to the specific flight situation. The observer pays attention to the battery status, flight speed, flight altitude, flight attitude, and route completion status displayed in the flight control software. After completing the crack image shooting task, the drone is landed at the predetermined location.
[0095] Furthermore, the dataset module includes, 1) image inspection and preprocessing, importing the images taken by the drone into the local device, checking whether the images taken by the drone are blurry, using the blur denoising method to process the blurrier images, and if the quality of the processed images is still low, they can be excluded. Then use cropping, color dithering, and scaling to expand the data. 2) Dataset construction, manually annotating the images taken by the drone and dividing them into two categories: images with cracks and images without cracks. Merge the crack images publicly available on the Internet with the crack images taken by the drone, convert all the images into RGB images in PNG format, and put them into two folders, one with cracks and one without cracks, to complete the dataset construction.
[0096] Furthermore, the training module includes a lightweight Transformer network and its training and testing, wherein the lightweight Transformer includes three parts: input conversion, Transformer encoder and prediction head, wherein the input conversion includes blocking, mixed block conversion and linear projection, the Transformer encoder mainly includes two parts: a multi-head attention layer and a feedforward network, and the prediction head includes a global average pooling layer and a linear layer; position encoding needs to be inserted between the input conversion and the Transformer encoder; the mixed block conversion method discards part of the Transformer input sequence during training, forcing the Transformer to use an incomplete sequence for training, and this method is not used during testing, and the complete input sequence is used for calculation.
[0097] Furthermore, the detection module includes a deployment unit for deploying the trained lightweight Transformer model to the AI development board; an unloading unit for performing calculation unloading; a calculation unit for transmitting the unloaded bridge concrete crack real-time detection task to the edge computing device for calculation; and a result unit for transmitting the calculation result back to the ground-end device. The NVIDIA Jetson TX2 development board is used as the deployment platform for the Transformer model, and Jetpack is used to install the operating system and necessary deep learning libraries for the NVIDIA Jetson TX2. The lightweight Transformer Pb model trained by TensorFlow is converted into an uff model available for TensorRT, and the convert script provided by the uff package is called to complete the model conversion. The converted uff model is deployed using TensorRT to generate the corresponding TensorRT inference engine, and then the input prediction image can be inferred. The NVIDIA Jetson TX2 development board is mounted on the drone as an edge computing device. During the flight of the drone, the camera system it carries will obtain the video stream image data of the bridge surface, and the unloaded data stream will be transmitted to the edge computing device through the formulated calculation unloading scheme. Then, the AI development board on the drone acts as an edge computing server, using the lightweight Transformer model deployed in it for detection, and transmits the recognition result and confidence level of whether the video frame image is a crack back to the ground station. Finally, the ground-end device displays the detected category and confidence level in real time in the received video stream. Through the cooperation of the ground-end program, the drone flight control program, and the airborne edge computing program, the drone can inspect the set area and detect AI bridge concrete cracks.
[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A real-time detection method for bridge concrete cracks based on edge computing and Transformer, characterized in that: The steps include: Step S1: Collecting bridge concrete crack images; including collecting public crack images and using drones to take crack images; Step S2: Check and preprocess the images to establish a data set; check the images taken by the drone, and use the fuzzy denoising method to process the blurry images; use cropping, color dithering, and scaling to expand the data; manually annotate the images taken by the drone and divide them into two categories: images with cracks and images without cracks; merge the crack images publicly available on the Internet with the crack images taken by the drone to construct a data set; Step S3: Construct a lightweight Transformer network suitable for edge computing; the network includes three parts: input conversion, Transformer encoder and prediction head, wherein the input conversion includes block division, mixed block transformation and linear projection, the image needs to be divided into 8*8 image blocks before inputting the network, and then the mixed block transformation is performed, and the input sequence of the Transformer encoder is formed after linear projection, and the embedding dimension of each block is 128; position coding needs to be inserted between the input conversion and the Transformer encoder; the Transformer encoder mainly includes two parts: a multi-head attention layer and a feedforward network FFN, wherein the multi-head attention layer has 2 heads, the first layer dimension of FFN is 512, the second layer dimension is 128, and the activation function is GELU; the prediction head includes a global average pooling layer and a linear layer, wherein the linear layer is a two-layer perceptron, the input dimension is 128, the first layer output dimension is 512, the second layer is used for classification, and the output dimension is the number of categories 2, representing crack-free images and cracked images respectively; The hybrid block transform method includes two types of transform methods, as follows: The first type of transformation is block discarding transformation, in which half of the non-adjacent blocks are discarded from all the blocks obtained after segmentation. Let P be the block matrix after segmentation, and the main formula of the block discarding method is as follows: P=P*D Where P has i rows and j columns, and D is a binary matrix (0 or 1), where 0 represents the block to be discarded and 1 represents the block to be retained; the binary matrix D is calculated by the following formula: D = CycleAndReshape(0,1,n) First, a binary array is generated, which is composed of alternating cycles of 0 and 1. n = i*j represents the total number of image blocks and is the length of the generated binary array. Then, the binary array is reconstructed into a matrix of the same size as P. Finally, the P matrix is multiplied by the D matrix, and the resulting matrix is the transformed block matrix. The second type of transformation is block thumbnail transformation, which randomly selects half of all image blocks and converts them into thumbnails, with the size converted from the original 8*8 to 4*4; these thumbnails are mixed with the other half of the image blocks in pairs, and only the mixed image blocks are retained. There are two sample blocks patch1 and patch2, and the mixing formula is as follows: patch mix =Mpatch1+Fill(Tn(patch2)) Among them, patch mix is a new mixed sample block, M is a binary mask representing the area where the image block is cropped and retained, the Fill function fills and generates an image block of the same size as patch2, and the Tn function is a thumbnail operation; The hybrid block transform is a method used in the model training stage, which transforms the input image blocks during training, discards part of the input information, and allows the Transformer to use an incomplete input sequence for training, while using the complete input sequence for calculation during inference; during training, the hybrid block transform uses a transformation method to transform the input image blocks each time, and the transformation method used can be controlled by a manually set probability threshold according to actual needs; Step S4: training the network to obtain a recognition model; pre-training the lightweight Transformer network on the large-scale image dataset ImageNet-1K, and then fine-tuning the pre-trained model on the bridge concrete crack dataset obtained in step S2; Step S5: Deploy the recognition model to the drone and use the drone to perform real-time detection of bridge concrete cracks; deploy the trained lightweight Transformer model to the AI development board; formulate a calculation offloading plan based on the model's computing power, transmission latency and other information; mount the AI development board on the drone as an edge computing device, and when the drone performs real-time detection, it will transfer the offloaded bridge concrete crack detection task to the edge computing device for calculation, and transmit the identified category and confidence back to the ground device.
2. A real-time detection device for bridge concrete cracks based on lightweight Transformer, characterized in that: Includes the following modules: An acquisition module, which acquires images of bridge concrete cracks; The data set module performs data inspection and preprocessing on the images and establishes the data set; Network module, building a lightweight Transformer network suitable for edge computing; the lightweight Transformer includes three parts: input conversion, Transformer encoder and prediction head, wherein the input conversion includes block, mixed block transformation and linear projection, the Transformer encoder mainly includes two parts: multi-head attention layer and feedforward network, and the prediction head includes global average pooling layer and linear layer; position encoding needs to be inserted between the input conversion and the Transformer encoder; The hybrid block transformation method includes: the first type of transformation is block discarding transformation, in which half of the non-adjacent blocks are discarded from all the blocks obtained after segmentation; the second type of transformation is block thumbnail transformation, in which half of all the image blocks are randomly selected and converted into thumbnails, and the size is converted from the original 8*8 to 4*4; these thumbnails are mixed with the other half of the image blocks in pairs, and only the mixed image blocks are retained; the hybrid block transformation is a method used in the model training stage, which transforms the input image blocks during training, discards part of the input information, and allows the Transformer to use an incomplete input sequence for training, and uses a complete input sequence for calculation during inference; during training, the hybrid block transformation uses a transformation method to transform the input image blocks each time, and the transformation method used can be controlled by a manually set probability threshold according to actual needs; Training module, training the network to obtain the recognition model; The detection module deploys the recognition model to the drone and uses the drone to perform real-time detection of bridge concrete cracks. It includes a deployment unit, which is used to deploy the trained lightweight Transformer model to the AI development board; an unloading unit, which is used to offload calculations; a computing unit, which is used to transmit the offloaded real-time detection of bridge concrete cracks tasks to the edge computing device for calculation; and an output unit, which is used to transmit the calculation results back to the ground-end device.