Deep learning-based bridge floor data cleaning method and system

By introducing the MultKAN module of multiplication and ReLU functions into the deep learning model, combined with the Resnet and Transformer frameworks, the problem that traditional methods are difficult to quickly and accurately classify bridge pavement data is solved, and efficient and accurate data cleaning and classification are achieved.

CN119992127AActive Publication Date: 2025-05-13EAST CHINA JIAOTONG UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411953838.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-13
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

Traditional manual processing methods are difficult to quickly and accurately clean and classify the database of bridge pavement, resulting in high resource consumption, poor timeliness and inaccurate classification results.

Method used

The Resnet-MultKAN-Transformer model based on deep learning is adopted, combining multiplication and ReLU functions to build a model for bridge pavement data cleaning, and images are taken by drones and databases are constructed for training and classification.

Benefits of technology

It realizes the rapid and accurate cleaning and classification of bridge pavement data, which is suitable for large data sets, reduces resource consumption, improves classification speed and accuracy, and ensures the accuracy of the cleaning and classification results of bridge pavement database.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992127A_ABST
    Figure CN119992127A_ABST
Patent Text Reader

Abstract

The invention relates to a bridge floor data cleaning method and system based on deep learning, the cleaning method uses a Resnet-MultKAN-Transformer model based on multiplication and ReLU function to train data in a database of a bridge road surface, and the cleaning method comprises the following specific steps: step 1, using an unmanned aerial vehicle carrying a high definition camera to shoot the road surface of the bridge, and using the unmanned aerial vehicle carrying the high definition camera to shoot the road surface of the bridge; the shot images are divided into three types: a bridge floor of a bridge, other structures of the bridge except the bridge floor structure, and a non-bridge structure; building a database of the bridge pavement by using all the shot images; step 2, constructing a Resnet-MultKAN-Transform model based on a multiplication function and a ReLU function, and adopting a framework of the Resnet model and the Transform model; and step 3, training a model by using the database constructed in the step 1, and using the trained model for cleaning and classifying bridge floor data. A MultiKAN module with multiplication and activation functions is added in a Resnet structure and a Transform, so that the problems of inaccurate data cleaning and classification, poor timeliness and huge workload of a traditional data processing method on a bridge pavement can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bridge deck data cleaning, and in particular to a bridge deck data cleaning method and system based on deep learning. Background Art

[0002] With the rapid development of modern economy, bridges have become important transportation infrastructure, and a large amount of data is generated during their design, construction, operation and maintenance. The content of these data includes but is not limited to the pavement of the bridge, as well as other bridge structures and some image data unrelated to the bridge. Due to the huge volume, diversity and complexity of these data, traditional manual processing methods often find it difficult to quickly screen the parts of the bridge, which requires a lot of manpower and material resources. Therefore, it is of great significance to clean and classify the database of bridge pavement.

[0003] Since traditional manual processing methods are difficult to process data in bridge and pavement databases, they have a great impact on bridge and pavement defect detection. However, with the continuous maturity of deep learning network model technology, the cleaning and classification of bridge and pavement databases are developing in this direction. Therefore, combining the deep learning-based network model with bridge and pavement data cleaning to quickly and accurately realize the cleaning and classification of bridge deck data is a technical problem that needs to be solved urgently. Summary of the invention

[0004] The purpose of the present invention is to provide a bridge deck data cleaning method and system based on deep learning, which can quickly and accurately clean the bridge pavement data and perform classification.

[0005] To achieve the above object, the technical solution of the present invention is:

[0006] In a first aspect, the present invention provides a bridge deck data cleaning method based on deep learning, wherein the cleaning method uses a Resnet-MultKAN-Transformer model based on multiplication and ReLU functions to train data in a database of bridge pavement, and the specific steps are:

[0007] Step 1: Use a drone with a high-definition camera to shoot the pavement of the bridge. The position of the drone changes constantly during the bridge construction process. The drone shoots in the direction of the front height of the object being photographed. The captured images contain three categories: the first is the bridge deck, the second is other bridge structures except the bridge deck structure, and the third is non-bridge structures. Use all the images taken to build a database of bridge pavement;

[0008] Step 2: Build a Resnet-MultKAN-Transformer model based on multiplication and ReLU functions:

[0009] The Resnet-MultKAN-Transformer model based on multiplication and ReLU functions adopts the framework of Resnet and Transformer models, including 7×7 convolutional layers, 3×3 maximum pooling downsampling layers, 3 ReLU-MultKAN modules, 6 3×3 convolutional layers, ReLU-MultKAN-Transformer converters, average pooling layers and fully connected layers;

[0010] The ReLU-MultKAN-Transformer converter includes a position encoding module, a ReLU-MultKAN-Transformer encoder and a ReLU-MultKAN-Transformer decoder, both of which are provided with a ReLU-MultKAN module;

[0011] The ReLU-MultKAN module includes a 1×1 convolutional layer, a batch normalization layer, a ReLU-MultKAN structure, a 3×3 convolutional layer, a batch normalization layer and a 1×1 convolutional layer connected in sequence;

[0012] In the ReLU-MultKAN structure, each layer has a series of nodes, some of which are multiplied by the activation function ω(x), while others are transformed by the same transformation. The activation function ω(x) is the product of the basic function b(x) and the function R i The sum of (x) is expressed as:

[0013] ω(x)=w(b(x)+R i (x)),

[0014] Wherein, w is used to control the overall size of the activation function; the basis function b(x) is a well-known part in the art;

[0015] Function R i (x) is expressed as:

[0016]

[0017] Where ReLU(x)=max(0,x); s i and e i is the domain of the independent variable x of the function to be fitted, that is, x∈[s i , e i ];

[0018] Step 3: Use the database constructed in step 1 to train the Resnet-MultKAN-Transformer model based on multiplication and ReLU functions, and use the trained Resnet-MultKAN-Transformer model based on multiplication and ReLU functions for cleaning and classification of bridge deck data.

[0019] Furthermore, the cleaning and classification results are divided into three types of images, namely, the bridge deck structure, other bridge structures other than the bridge deck structure, and non-bridge structures; and according to the classified data images, they are classified into positive samples and negative samples;

[0020] Data images of bridge deck structures classified as bridges are positive samples;

[0021] Data images classified as other bridge structures other than bridge deck structures and non-bridge structures are negative samples.

[0022] Furthermore, during the training process, the learning rate is set to: The formula for the change of the learning rate of the i-th generation is:

[0023] Lr=V min +(LV min )*(1+cos(π*Le / T max )) / 2.

[0024] Where Lr is the result of the learning rate; T max is the total number of training cycles, used to define the period of the cosine function; V min is the minimum value of the learning rate; L e is the current number of training cycles; L is the basic learning rate;

[0025] The loss function is the weighted cross entropy loss function Loss, and the formula is:

[0026]

[0027] Where z is the true distribution of each pixel, is the output distribution of the model.

[0028] Furthermore, in the ReLU-MultKAN-Transformer encoder, the feature sequence is input into the multi-head mask layer together with the output of the position encoding module, and then input into the ReLU-MultKAN module. After passing through the ReLU-MultKAN module, the result is residually connected with the feature sequence input into the ReLU-MultKAN-Transformer encoder, and the result is output to the normalization layer of the lower layer, and then input into the ReLU-MultKAN feedforward network module. Then, the output of the ReLU-MultKAN feedforward network module is residually connected with the input, and then the result of the residual connection is input into the normalization layer to obtain the output of the ReLU-MultKAN-Transformer encoder.

[0029] The output of the ReLU-MultKAN-Transformer encoder is input into the first multi-head mask layer in the ReLU-MultKAN-Transformer decoder together with the positional encoding output by the positional encoding module, and then input into the ReLU-MultKAN module of the lower layer. The output of the ReLU-MultKAN module is residually connected with the output of the ReLU-MultKAN-Transformer encoder and then input into the normalization layer. The output of the normalization layer is then input into the second multi-head mask layer together with the positional encoding, and then input into the ReLU-MultKAN module. Its output is residually connected with the feature sequence of the second multi-head mask layer, and then input into the normalization layer of the lower layer, and then input into the ReLU-MultKAN feedforward network module. The output of the ReLU-MultKAN feedforward network module is residually connected with the input, and the result is output into the normalization layer to obtain the output of the ReLU-MultKAN-Transformer converter.

[0030] Furthermore, the ReLU-MultKAN feed-forward network module includes a reshaping and upgrading module, a ReLU-MultKAN module and a CBL module.

[0031] In a second aspect, the present invention provides a bridge deck data cleaning system based on deep learning, and the system executes the steps.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] The present invention creatively adds a MultKAN module with multiplication and activation functions to the Resnet structure and Transformer in the Resnet-MultKAN-Transformer model, and constructs a Resnet-MultKAN-Transformer model based on multiplication and ReLU functions, which can solve the problems of inaccurate data cleaning and classification of bridge pavement, poor timeliness and huge workload of traditional data processing methods.

[0034] The cleaning method of the present invention is used for cleaning and classifying bridge and pavement data images, and is particularly suitable for cleaning and classifying image data of large data sets (with tens of thousands or hundreds of thousands of images). The effect is more obvious. For the processing of large amounts of data in the bridge engineering construction cycle, the resource consumption is reduced, the classification speed is accelerated, and the accuracy of the classification results is improved. Moreover, the cleaning and classification problems of the bridge and pavement database can be completed accurately and efficiently, providing accurate guarantee for the subsequent use of the bridge and pavement database.

[0035] The present invention defines a ReLU function to improve the MultKAN structure, which speeds up the progress of model training, reduces the complexity of cleaning and classification of the bridge pavement database, reduces the amount of data, and ensures the accuracy of data for defect detection and analysis of the bridge pavement. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a structural diagram of a Resnet-MultKAN-Transformer model based on multiplication and ReLU functions in an embodiment of the present invention.

[0037] Figure 2 It is a schematic diagram of the structure of the ReLU-MultKAN-Transformer converter in the present invention.

[0038] Figure 3 It is a structural schematic diagram of the ReLU-MultKAN structure in the present invention.

[0039] Figure 4 It is a structural diagram of the ReLU-MultKAN module in the present invention.

[0040] Figure 5 It is a schematic diagram of the structure of the ReLU-MultKAN feed-forward network module (ReLU-MultKAN-FFN) in the present invention. DETAILED DESCRIPTION

[0041] The present invention is further explained below in conjunction with the embodiments and drawings, but this is not intended to limit the scope of protection of the present application.

[0042] Example 1

[0043] The bridge deck data cleaning method based on deep learning in this embodiment is: obtaining data from a database of bridge pavement, using a Resnet-MultKAN-Transformer model based on multiplication and ReLU functions to train the data in the database of bridge pavement to obtain appropriate parameter information, and using the trained model to perform cleaning and classification to obtain classification results. The method specifically includes the following steps:

[0044] Step 1: Get the data in the database of bridge pavement:

[0045] A drone equipped with a high-definition camera is used to photograph the road surface of the bridge. There is a certain distance between the drone and the road surface of the bridge being photographed. The position of the drone is constantly changing during the bridge construction process. The drone shoots in the direction of the front height of the object being photographed. The captured images will contain three categories. The first is the bridge deck, the second is other bridge structures except the bridge deck structure, and the third is non-bridge structures.

[0046] The camera in this embodiment uses a high-precision camera. The distance between the drone and the road surface of the bridge to be photographed is 20 meters. The bridge deck, other structures and non-bridge structures can be photographed. The drone will take pictures every 10 minutes and store the pictures in the memory. 48 pictures are taken every day. The construction period of the bridge in this embodiment is about 3 months, and a total of 12,960 pictures can be obtained, and the images taken are all PNG pictures of 512×512 pixels.

[0047] The captured images were randomly shuffled and divided into a training set and a validation set in a ratio of 8:2, with the number of images in each set being 10368 and 2592 respectively.

[0048] Step 2: Build a Resnet-MultKAN-Transformer model based on multiplication and ReLU functions:

[0049] The Resnet-MultKAN-Transformer model based on multiplication and ReLU functions adopts the framework of Resnet and Transformer models. The overall structure is as follows Figure 1 As shown, it includes 14 modules, namely 7×7 convolution layer, 3×3 maximum pooling downsampling layer, 3 ReLU-MultKAN modules, 6 3×3 convolution layers, ReLU-Multiplication Kolmogorov-Arnold Transformer (ReLU-MultKAN-Transformer Transformer), average pooling layer, and fully connected layer.

[0050] The ReLU-MultKAN module mainly includes a 1×1 convolution layer, a batch normalization layer, a ReLU-MultKAN structure, a 3×3 convolution layer, a batch normalization layer and a 1×1 convolution layer.

[0051] The image is input into the 7×7 convolutional layer of the first layer, then into the 3×3 maximum pooling downsampling layer of the second layer, and then into the ReLU-MultKAN module of the third layer.

[0052] In the ReLU-MultKAN module of the third layer, the feature sequence is first input into the 1×1 convolution layer, then into the batch normalization layer, followed by the ReLU-MultKAN structure, the 3×3 convolution layer, the batch normalization layer and the 1×1 convolution layer. Then the output of the 1×1 convolution layer is residually connected to the feature sequence of the input ReLU-MultKAN module, and then the result is output to the 3×3 convolution layer of the fourth layer. Then its output is output to the 3×3 convolution layer of the fifth layer. Its output is residually connected to the feature sequence of the input ReLU-MultKAN module of the third layer, and then input to the ReLU-MultKAN module of the sixth layer. The output of the eLU-MultKAN module is input into the 3×3 convolutional layer of the seventh layer, and then the output of this layer is input into the 3×3 convolutional layer of the eighth layer, and then its output is residually connected with the feature sequence of the ReLU-MultKAN module input into the sixth layer, and the result of the connection is input into the ReLU-MultKAN module of the ninth layer, and the feature sequence of this module is input into the 3×3 convolutional layer of the tenth layer, and then its output is input into the 3×3 convolutional layer of the eleventh layer, and then its output is residually connected with the input of the ReLU-MultKAN module of the ninth layer, and the result of the connection is input into the ReLU-Multiplication Kolmogorov-Arnold Converter of the twelfth layer.

[0053] The ReLU-MultKAN-Transformer converter includes a position encoding module, a ReLU-MultKAN-Transformer encoder and a ReLU-MultKAN-Transformer decoder.

[0054] In the ReLU-MultKAN-Transformer encoder, the feature sequence is input into the multi-head mask layer together with the output of the position encoding module, and then input into the ReLU-MultKAN module. After passing through the ReLU-MultKAN module, the result is residually connected with the feature sequence input into the ReLU-MultKAN-Transformer encoder, and the result is output to the normalization layer of the lower layer, and then input into the ReLU-MultKAN feedforward network module. The output and input of the ReLU-MultKAN feedforward network module are residually connected, and the result of the residual connection is input into the normalization layer to obtain the output of the ReLU-MultKAN-Transformer encoder.

[0055] The output of the ReLU-MultKAN-Transformer encoder is input into the first multi-head mask layer in the ReLU-MultKAN-Transformer decoder together with the positional encoding output by the positional encoding module (the positional encoding module here is the same as the positional encoding module in the encoder module), and then input into the ReLU-MultKAN module in the lower layer. The output of the ReLU-MultKAN module is residually connected with the output of the ReLU-MultKAN-Transformer encoder, and the result is input into the normalization layer, and then The output of the normalization layer and the position code are input into the second multi-head mask layer together, and then into the ReLU-MultKAN module. The output and the feature sequence of the second multi-head mask layer are residually connected, and the result is input into the normalization layer of the lower layer, and then into the ReLU-MultKAN feed-forward network module. The output and input of the ReLU-MultKAN feed-forward network module are residually connected, and the result is output into the normalization layer, and then into the average pooling layer of the thirteenth layer. Finally, the output of the average pooling layer of the thirteenth layer is input into the fully connected layer of the fourteenth layer, and finally the classification result is output.

[0056] Step 3: Use the database constructed in step 1 to train the Resnet-MultKAN-Transformer model based on multiplication and ReLU functions, and use the trained Resnet-MultKAN-Transformer model based on multiplication and ReLU functions for cleaning and classification of bridge deck data.

[0057] First, the hyperparameters (initial weights, learning rates, etc.) of the Resnet-MultKAN-Transformer model based on multiplication and ReLU functions are randomly initialized, and the model is trained for 50 epochs using the training set. After each epoch of training, the validation set is used to verify the generalization ability of the deep learning model and prevent overfitting.

[0058] The learning rate is set as follows: the cosine decay strategy is adopted, and the smooth adjustment strategy is used during the training process, that is, the learning rate is gradually adjusted, which can improve the convergence and performance of the model. The formula for the change of the i-th generation learning rate is as follows:

[0059] Lr=V min +(LV min )*(1+cos(π*Le / T max )) / 2.

[0060] Where Lr is the result of the learning rate. max is the total number of training epochs, which is used to define the period of the cosine function. max =50, V min is the minimum value of the learning rate. During the training process, the learning rate will not be lower than this value. Le is the current number of training cycles. L is a basic learning rate.

[0061] Secondly, the loss function is used to adjust the parameter model. The loss function of the model of the present invention is a weighted cross entropy loss function Loss, and the formula is as follows:

[0062]

[0063] Where z is the true distribution of each pixel, is the output distribution of the model.

[0064] Set up a visualization chart to display the learning rate and loss function curves in real time. During the training process, the current accuracy and time consumption will be output after each epoch. By observing the trend of the curve, analyze whether the convergence and accuracy of the model meet the expected requirements. If the model cannot converge or the accuracy is low after convergence, it needs to be adjusted by adjusting the hyperparameters. When the value of the loss function Loss is less than 0.001, it means that the model has reached convergence and the training is stopped.

[0065] The data images in the database of the classified bridge pavement to be cleaned are input into the trained Resnet-MultKAN-Transformer model based on multiplication and ReLU functions for recognition, and three types of images are obtained, namely, the bridge deck structure, other bridge structures other than the bridge deck structure, and non-bridge structures. And according to the classified data images, they are classified into positive samples and negative samples.

[0066] The data images of the bridge deck structure classified as a bridge are positive samples and can be used as data in the database for defect detection of the bridge deck.

[0067] Data images classified as other bridge structures and non-bridge structures other than bridge deck structures are negative samples and are needed for other detections.

[0068] Through the data cleaning method of the present invention, the collected image data can be quickly cleaned and classified, and the classified data can be used for the establishment and update of the bridge pavement data set in the later stage. After this classification processing, different samples can be labeled to facilitate the later maintenance, verification and inspection, etc.

[0069] Example 2

[0070] The structures of the various parts of the Resnet-MultKAN-Transformer model based on multiplication and ReLU functions in this embodiment are:

[0071] First, the image I0 is input into the 7×7 convolutional block of the first layer. After the convolution operation of this module, it is then passed through the 3×3 maximum pooling downsampling layer of the second layer. After the 3×3 maximum pooling downsampling layer, the output feature sequence I1 is 1 / 4 of the original image. Then it is input into the ReLU-MultKAN module of the third layer.

[0072] Next, the feature sequence of size 1 / 4 is input into the ReLU-MultKAN module of the third layer. In the ReLU-MultKAN module, it is first input into the 1×1 convolution layer of the first layer. After the convolution process, it is then input into the batch normalization layer, which calculates the mean and variance of the output of the upper layer and uses these statistics to normalize the output of the layer. By reducing the internal covariate shift, the model can stabilize the training process and converge to the optimal solution faster in the process. Then the ReLU-MultKAN structure of the lower layer is input. Each layer of the structure will have a series of nodes, and there will be a set of specific functions in the nodes to process the input data and output it to the next layer.

[0073] The standard KAN layer Φ is expressed as:

[0074]

[0075] Among them, φ l,1,1 (·) represents a one-dimensional function, l represents the depth, n l belongs to the integer array [n1, n2, ..., n l ], x l represents the input vector, and The th KAN layer has n l Input dimensions and n l+1 Output dimension, the input vector x l Convert to x l+1 .

[0076] The entire KAN(x) network is composed of standard KAN layers, namely

[0077]

[0078] Among them, Φ l is a standard KAN layer. The number of addition and multiplication operations in layer l is expressed as and in Φ l Accepts an input vector And Φ l Convert to output z l ,Right now

[0079]

[0080] Multiplication layer M l Including identity transformation operations and multiplication operations, for some child nodes z l Use the activation function ω(x) to perform a multiplication operation (such as Figure 3 The black multiplication sign in the figure shows that the same transformation operation is performed on the other child nodes ( Figure 3 The black solid circle in the middle layer). The multiplication layer M of the ReLU-MultKAN structure l The activation function ω(x) of the node in the multiplication operation is the basis function b(x) and the function R i The sum of (x), ω(x) can be expressed as:

[0081] ω(x)=w(b(x)+R i (x)),

[0082] Among them, w is used to control the overall size of the activation function. The basic function b(x) is expressed as:

[0083]

[0084] Function Ri (x) is expressed as:

[0085]

[0086] Where ReLU(x)=max(0,x). i and e i is the domain of the independent variable x of the function to be fitted, that is, x∈[s i , e i ].

[0087] Then the ReLU-MultKAN layer represents the multiplication layer M l With the standard KAN layer Φ l The operation can be concisely expressed as:

[0088]

[0089] The entire ReLU-MultKAN structure is expressed as:

[0090]

[0091] The present invention introduces the ReLU-MultKAN module into the model, which can achieve performance comparable to or even better than that of the traditional convolutional network with fewer parameters, has a small memory requirement, and has a learnable nonlinear activation function. It can significantly improve the computational efficiency and can be flexibly adapted to different application scenarios. Here, the original B-spline function is replaced by the function R i (x), so that the model can make full use of the parallel capabilities of the GPU, reduce the time of model training, and reduce resource consumption. The ReLU-MultKAN structure can fit the nonlinear transformation that best fits the data at each degree. Each input variable is embedded through a set of independent one-dimensional nonlinear learnable functions φ p,q (x i ), where p represents the index of the input variable and q represents the index of the output dimension. The ReLU-MultKAN structure is based on the Kolmogorov-Arnold representation theorem, which states that a multivariate continuous function can be represented as a finite combination of a single variable continuous function, which can be expressed as:

[0092]

[0093] Among them, φ q,p (x p) is a one-dimensional function, p is the index of the input dimension, q is the index of the output dimension, and n is the number of parameters referenced. The activation function within the network node in the ReLU-MultKAN structure avoids errors in sensitive data input, thereby reducing the model's convolution and pooling operations on sensitive data, so that the ReLU-MultKAN structure further improves the accuracy of the output image and allows the ReLU-MultKAN structure to achieve complex multi-dimensional input to multi-dimensional output mapping at each layer.

[0094] After passing through the ReLU-MultKAN structure, the feature sequence will be input into the 3×3 convolutional layer of the lower layer. After the convolution operation, it will be input into the batch normalization layer of the lower layer. After the batch normalization layer, it will be input into the 1×1 convolutional layer. Its output will be residually connected to the input of the ReLU-MultKAN module. The residual connection result is the output of the ReLU-MultKAN module.

[0095] Then the output of the ReLU-MultKAN module of the third layer is input into the 3×3 convolution layer of the fourth layer, and the output of this layer is input into the 3×3 convolution layer of the fifth layer. After the convolution operation of this layer, the feature sequence at this time is I2, and its size is 1 / 8 of the original image. Then the feature sequence is residually connected with the input of the ReLU-MultKAN module of the third layer, and the connected result is output to the ReLU-MultKAN module of the sixth layer, and then input into the 3×3 convolution layers of the seventh and eighth layers. After the eighth 3×3 convolution layer, the feature sequence is residually connected with the feature sequence I2 input into the ReLU-MultKAN module of the sixth layer. The result of the connection is I3, and the size is 1 / 16 of the original image. Then I3 is input into the ReLU-MultKAN module of the ninth layer, and then into the 3×3 convolution layer of the tenth layer and the 3×3 convolution layer of the eleventh layer. After the feature sequence is convolved, its output will be residually connected with the feature sequence I3. The result of the connection is I4, which is 1 / 32 of the original image. The feature sequence I4 is input into the ReLU-MultKAN-Transformer converter of the twelfth layer.

[0096] Before the feature sequence I4 enters the ReLU-multiplication Kolmogorov-Arnold encoder, it will first perform an Embedding operation, that is, mapping the feature sequence I4 into a vector form. The Embedding operation is divided into two parts. The first part is the Input Embedding operation, which maps the feature sequence I4 into a vector a. At this time, the size of vector a is 1 / 32 of the original image. The second part is the Positional Encoding operation, which is a set of vectors b with the same dimension as the vector after the Input Embedding operation, which is used to provide position information. Regarding the positional encoding operation, its encoding rules are as follows:

[0097]

[0098]

[0099] Among them, PE represents the position encoder, pos x,y Indicates pos x or pos y , (pos x ,pos y ) is a position in the feature sequence, and i represents the dimension of the position. x Substituting the two encoding formulas, we can calculate a 16-dimensional vector representing pos x Position encoding; pos y Substituting the last two encoding formulas, we can also calculate a 16-dimensional vector, representing pos y Position encoding; by concatenating these two vectors, we can get a vector with the same dimension as the input feature sequence, representing (pos x ,pos y ) position encoding. By calculating all the position encodings, we get a vector with the same dimension as the input feature sequence, which represents the position encoding of this batch.

[0100] Next, we need to add vector a and vector b to get vector c, which is the key K and query Q in the multi-head mask layer in the ReLU-MultKAN-Transformer encoder. The multi-head mask layer can focus on different parts of the input vector c at the same time, and map the query, key, and value to multiple different linear spaces respectively. Vector a is also used as the value V of the multi-head mask layer. For each set of inputs (query, key, value) of vector c and vector a, four steps are performed. First, the query Q, key K, and value V are obtained, which can be expressed as:

[0101] Q=W q ×query,

[0102] K=W k ×key,

[0103] V=W v ×value.

[0104] Among them, W q , W k and W v is the weight matrix, corresponding to different heads.

[0105] The second step is for each head to calculate the similarity between the query and the key and use it as the attention score. It can be expressed as:

[0106]

[0107] Among them, Attention_scores[i] represents the attention score of the i-th position in the feature sequence, softmax is the function, Q i represents the query vector at the i-th position in the feature sequence, K i represents the key vector of the i-th position in the feature sequence, d_k represents the dimension of the key vector, and the dimension under the square root is divided here to stabilize the gradient and prevent the value from being too large.

[0108] The third step is to apply the attention score to the Value. Each head has its own score, so a sum is taken for each head, which can be expressed as:

[0109] Attention_output[i]=Attention_scores[i]×V,

[0110] Among them, Attention_output[i] represents the weighted summation result of the i-th position, and V represents the value Value.

[0111] Finally, the results of all heads are concatenated and integrated into the final output vector through another linear transformation. The multi-head mask layer mechanism can capture richer dependencies. Each head processes part of the information of the input vector c and vector a. Combining the results of all heads can enhance the understanding of the global context. The final output result and the input vector a are combined into the feature sequence I5 and output to the ReLU-MultKAN module.

[0112] In the ReLU-MultKAN module, the feature sequence I5 needs to pass through the 1×1 convolution layer, batch normalization layer, ReLU-MultKAN structure, 3×3 convolution layer, batch normalization layer and 1×1 convolution layer in sequence, and perform residual connection with the feature sequence I5 at the output. The result of the connection will be residually connected with the feature sequence I4, and the result of the connection will be input into the normalization layer of the lower layer. In the normalization layer, the feature sequence can avoid the disappearance of the gradient. At the same time, this layer improves the stability and generalization ability of the model. The output I6 of the normalization layer will be input into the ReLU-MultKAN feedforward network module of the lower layer.

[0113] The ReLU-MultKAN feedforward network module includes a reshaping and upgrading module, a ReLU-MultKAN module, and a CBL module. First, the reshaping and upgrading module reshapes the feature sequence I6 and refines it. After that, it passes through a ReLU-MultKAN module, performs further convolution operations on the feature sequence I6, and finally enters the CBL module. The CBL module contains 3 convolution blocks, where the convolution kernel size of the first two convolution blocks is 3×3, and the convolution kernel size of the third convolution block is 1×1. The CBL module is expressed in mathematical formula as follows:

[0114]

[0115] Among them, d1 is the output of the ReLU-MultKAN module in the ReLU-MultKAN feed-forward network module, is the output of the first convolutional block in the CBL module, is the output of the second convolutional block in the CBL module, C h1 and C h2 are the hidden layers of the first and second convolution blocks, G is the third convolution block, Leaky_ReLU is the Leaky_ReLU activation function, Batch is the batch normalization operation, and C b is the number of channels of feature sequence I6. The output of the ReLU-MultKAN feedforward network module is residually connected to feature sequence I6, and the result of the residual connection is input to the normalization layer of the lower layer. After the normalization layer, its output is feature sequence I7, which is then input to the ReLU-MultKAN-Transformer decoder of the lower layer.

[0116] Before the feature sequence I7 is input into the multi-head mask layer, an embedding operation is performed to map the output feature sequence I7 to a vector d, where the size of the vector d is 1 / 32 of the original image. This is followed by a position encoding operation to generate a vector e. Next, vector d and vector e need to be added to obtain vector f, which is the key K and query Q in the multi-head mask layer of the first layer in the ReLU-Kolmogorov-Arnold decoder. The multi-head mask layer maps the query, key, and value to multiple different linear spaces respectively. Vector d is also used as the value V of the multi-head mask layer. For each set of inputs (query, key, value) of vector f and vector d, the same four-step operation is performed as in the multi-head mask layer in the encoder. Finally, the output result and the input vector d are merged into the feature sequence I8 and output to the ReLU-MultKAN module. After the feature sequence I8 has undergone the 1×1 convolution layer, batch normalization layer, ReLU-MultKAN structure, 3×3 convolution layer, batch normalization layer and 1×1 convolution layer in the ReLU-MultKAN module, a new feature sequence is generated. The new feature sequence will be residually connected with the feature sequence I7. The result after the residual connection is input into the normalization layer of the lower layer. After the normalization operation, a new feature sequence I9 is ​​generated. Before the feature sequence I9 is ​​input into the multi-head mask layer of the lower layer, the embedding operation is performed to map the output feature sequence I9 to a vector g. At this time, the size of the vector g is 1 / 32 of the original image. Then the position encoding operation is performed to generate the vector h. Next, the vector g and the vector h need to be added to obtain the vector k, that is, the key K and the query Q in the multi-head mask layer of the fourth layer in the ReLU-multiplication Kolmogorov-Arnold decoder. The multi-head mask layer maps the query, key and value to multiple different linear spaces respectively. The vector g is also used as the value V of the multi-head mask layer. For each input set (query, key, value) of vector k and vector g, the same four-step operation is performed as the first multi-head mask layer in the decoder, and the final output result is merged with the input vector g into the feature sequence I 10 . Then the characteristic sequence I 10 The ReLU-MultKAN module will be input. After the operation of this module, a new feature sequence will be generated. This feature sequence will be residually connected with the feature series I9. The result of the residual connection will be input into the normalization layer of the lower layer for normalization operation to generate a new feature sequence I 11 . Then the characteristic sequence I 11 Input the ReLU-MultKAN feed-forward network module of the lower layer. The sequence will undergo the operations of the reshaping and upgrading module, the ReLU-MultKAN module and the CBL module to generate a new feature sequence. The sequence is finally input into the normalization layer of the lower layer for normalization operation to generate a new feature sequence I 12 .

[0117] Feature Sequence I 12 The input is sent to the average pooling layer, which averages the feature map of each channel into a value to obtain a one-dimensional vector with the same length as the number of channels. This layer can reduce the dimension of the feature map while retaining important feature information in the image. The one-dimensional vector output by the pooling layer is then input into the fully connected layer of the lower layer. The weight matrix of this layer maps the one-dimensional vector to the category space, that is, outputs the classified image.

[0118] The present invention adopts the framework of Resnet and Transformer, and uses the ReLU-MultKAN module therein, which can avoid the input of sensitive data and further improve the accuracy of the output classification image. The ReLU-MultKAN module is added to the traditional Resnet and Transformer model, which not only replaces the B-spline function in the MultKAN structural point with the ReLU activation function, so that the model can better and faster capture the relevant features in the image, speed up the training time of the model, and emphasize local attention, which helps to keep the model's attention to the feature details in the image, dynamically switch between local and global information, increase the model's adaptability to different features of the input image, help improve the model's extraction of features in complex images, enhance the model's ability to understand features in images, and also adjust the self-attention mechanism in the model, help the model to more effectively focus on the key areas in the image, improve the model's accuracy in extracting image features, and enhance the model's ability to classify images.

[0119] Table 1 below is a comparison of the effects of training different types of models using the embodiment data set, where F1 represents an indicator for evaluating the classification performance of the model, which is the harmonic mean of precision and recall. T refers to the time for model classification training.

[0120] Table 1

[0121]

[0122] The Resnet-MultKAN-Transformer model based on multiplication and ReLU functions of the present invention aims to solve the problems of long training time, large number of required parameters, high computational complexity, unintuitive feature fusion, strong data dependence, poor interpretability and high hardware requirements in the existing image classification models during image classification. By adding a method with a multiplication structure and a ReLU activation function in a suitable manner, the strong expression ability and interpretability of the ReLU-MultKAN structure can be combined with the feature extraction ability of Resnet and Transformer, thereby improving the accuracy of image classification. At the same time, the residual structure and self-attention mechanism of the two can be combined with the recognition-related feature ability of the ReLU-MultKAN structure, thereby optimizing the feature extraction and prediction process, reducing the model training time, improving the convergence speed of the model and the efficiency of input image classification, and improving the adaptability and precise fitting of the model, while improving the generalization ability of the model, and improving the accuracy of data cleaning classification of the bridge pavement database.

[0123] The present invention adopts the Resnet-MultKAN-Transformer model based on multiplication and ReLU functions, and adds the ReLU-MultKAN module, which greatly improves the accuracy of image feature extraction and classification, ensures the accuracy and timeliness of image feature extraction, and makes the model perform better in image classification tasks. Due to the innovative ReLU-MultKAN structure, the convergence speed and fitting ability of the model are significantly improved, which helps the model to achieve the best effect faster and can better capture the key features in the image, thereby improving the speed and accuracy of image classification. The database of bridge pavement can be cleaned and classified in a timely and accurate manner, which greatly reduces the time for data cleaning and classification of the bridge pavement database, and at the same time ensures the accuracy of the data for crack detection on the bridge pavement.

[0124] Any matters not described in the present invention are applicable to the prior art.

Claims

1. A bridge deck data cleaning method based on deep learning, characterized in that: The cleaning method uses the Resnet-MultKAN-Transformer model based on multiplication and ReLU functions to train the data in the database of bridge pavement. The specific steps are: Step 1: Use a drone with a high-definition camera to shoot the pavement of the bridge. The position of the drone changes constantly during the bridge construction process. The drone shoots in the direction of the front height of the object being photographed. The captured images are divided into three types: the first is the bridge deck, the second is other bridge structures except the bridge deck structure, and the third is non-bridge structures. Use all the captured images to build a database of bridge pavement; Step 2: Build a Resnet-MultKAN-Transformer model based on multiplication and ReLU functions: The Resnet-MultKAN-Transformer model based on multiplication and ReLU functions adopts the framework of Resnet and Transformer models, including 7×7 convolutional layers, 3×3 maximum pooling downsampling layers, 3 ReLU-MultKAN modules, 6 3×3 convolutional layers, ReLU-MultKAN-Transformer converters, average pooling layers and fully connected layers; The ReLU-MultKAN-Transformer converter includes a position encoding module, a ReLU-MultKAN-Transformer encoder and a ReLU-MultKAN-Transformer decoder, both of which are provided with a ReLU-MultKAN module; The ReLU-MultKAN module includes a 1×1 convolutional layer, a batch normalization layer, a ReLU-MultKAN structure, a 3×3 convolutional layer, a batch normalization layer and a 1×1 convolutional layer connected in sequence; In the ReLU-MultKAN structure, each layer has a series of nodes, some of which are multiplied by the activation function ω(x), while others are transformed by the same transformation. The activation function ω(x) is the product of the basic function b(x) and the function R i The sum of (x) is expressed as: ω x =w(b(x)+R i (x)), Among them, w is used to control the overall size of the activation function; Function R i (x) is expressed as: Where ReLU(x)=max(0,x); s i and e i is the domain of the independent variable x of the function to be fitted, that is, x∈[s i , e i ]; Step 3: Use the database constructed in step 1 to train the Resnet-MultKAN-Transformer model based on multiplication and ReLU functions, and use the trained Resnet-MultKAN-Transformer model based on multiplication and ReLU functions for cleaning and classification of bridge deck data.

2. The cleaning method according to claim 1, characterized in that: The cleaning and classification results are divided into three types of images, namely, the bridge deck structure, other bridge structures other than the bridge deck structure, and non-bridge structures; and according to the classified data images, they are classified into positive samples and negative samples; Data images of bridge deck structures classified as bridges are positive samples; Data images classified as other bridge structures other than bridge deck structures and non-bridge structures are negative samples.

3. The cleaning method according to claim 1, characterized in that: During the training process, the learning rate is set to: The formula for the change of the learning rate of the i-th generation is: Lr6V min +(LV min )*(1+cos(π*And / T max )) / 2。 Where Lr is the result of the learning rate; T max is the total number of training cycles, used to define the period of the cosine function; V min is the minimum value of the learning rate; Le is the current number of training cycles; L is the basic learning rate; The loss function is the weighted cross entropy loss function Loss, and the formula is: Where z is the true distribution of each pixel, is the output distribution of the model.

4. The cleaning method according to claim 1, characterized in that: In the ReLU-MultKAN-Transformer encoder, the feature sequence is input into the multi-head mask layer together with the output of the position encoding module, and then input into the ReLU-MultKAN module. After passing through the ReLU-MultKAN module, the result is residually connected with the feature sequence input into the ReLU-MultKAN-Transformer encoder, and the result is output to the normalization layer of the lower layer, and then input into the ReLU-MultKAN feedforward network module. The output of the ReLU-MultKAN feedforward network module is residually connected with the input, and the result of the residual connection is input into the normalization layer to obtain the output of the ReLU-MultKAN-Transformer encoder. The output of the ReLU-MultKAN-Transformer encoder is input into the first multi-head mask layer in the ReLU-MultKAN-Transformer decoder together with the positional encoding output by the positional encoding module, and then input into the ReLU-MultKAN module of the lower layer. The output of the ReLU-MultKAN module is residually connected with the output of the ReLU-MultKAN-Transformer encoder and then input into the normalization layer. The output of the normalization layer is then input into the second multi-head mask layer together with the positional encoding, and then input into the ReLU-MultKAN module. Its output is residually connected with the feature sequence of the second multi-head mask layer, and then input into the normalization layer of the lower layer, and then input into the ReLU-MultKAN feedforward network module. The output of the ReLU-MultKAN feedforward network module is residually connected with the input, and the result is output into the normalization layer to obtain the output of the ReLU-MultKAN-Transformer converter.

5. The cleaning method according to claim 4, characterized in that: The ReLU-MultKAN feed-forward network module includes a reshaping and upgrading module, a ReLU-MultKAN module and a CBL module.

6. A bridge deck data cleaning system based on deep learning, characterized in that: The system executes the steps described in any one of claims 1-5.

Citation Information

Patent Citations

  • Conv-KANformer-based neural network interference identification method

    CN118568590A

  • System and method of bridging the gap between object and image-level representations for open-vocabulary detection

    US20240203085A1