Prefabricated tunnel segment embedded sleeve detection method based on target detection algorithm

By applying a detection method based on the object detection algorithm in prefabricated tunnel pipe sheets, the problems of low efficiency and unstable accuracy of manual visual inspection are solved, and efficient and accurate detection of the number of embedded sleeves is achieved, which is suitable for complex environments.

CN119963804AActive Publication Date: 2025-05-09EAST CHINA JIAOTONG UNIVERSITY +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202411949247.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-09
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

The existing artificial visual inspection is in the detection of the number of embedded sleeves in prefabricated tunnel pipe sheets, and the accuracy is unstable, especially in environments such as obstruction of vision.

Method used

Using a detection method based on the object detection algorithm, the panoramic camera is used to obtain the tunnel pipe mold image, pre-process and feature extraction are performed through the KAN-DETR object detection model, combined with DIoU perceptual query selection and decoder to generate bounding box and confidence scores, and finally converted into specific quantities through the feedforward neural network for detection.

Benefits of technology

It greatly improves detection efficiency and accuracy, can quickly process large amounts of image data, automatically identify and count the embedded sleeves, reduce human error and missed inspection, and is suitable for detection needs in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963804A_ABST
    Figure CN119963804A_ABST
Patent Text Reader

Abstract

The invention discloses a prefabricated tunnel segment pre-embedded sleeve detection method based on a target detection algorithm. The method comprises the following steps that a panoramic camera is used for obtaining tunnel segment mold images where pre-embedded sleeves are installed and tunnel segment mold images where the pre-embedded sleeves are not installed before pouring; converting the size of the mold image into a standard size of 224 * 224, then multiplying each pixel by the same pixel weight, forming a data set by all new images, dividing the data set into a training set and a test set, and performing data annotation on the image containing the pre-embedded sleeve in the training set; carrying out mean filtering and data enhancement processing to form an enhanced data set; and constructing a KAN-DETR target detection model, and training the KAN-DETR target detection model by using the enhanced data set to detect the number of the pre-embedded sleeves of the prefabricated tunnel segment. The method has high discrimination speed and high discrimination precision, and is suitable for detecting the number of the pre-embedded sleeves.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of prefabricated tunnel segments, and in particular to a method for detecting embedded sleeves of prefabricated tunnel segments based on a target detection algorithm. Background Art

[0002] In the production of prefabricated tunnel segments, embedded components such as embedded channels, embedded steel plates and embedded sleeves play a key role in structural connection in prefabricated tunnel segments. They can ensure the tight connection between the segments, improve the stability and bearing capacity of the overall structure, and provide convenience for subsequent equipment installation and maintenance. Before pouring the tunnel segments, the detection of the number of embedded sleeves is usually based on manual visual inspection. Conventional manual visual inspection has the disadvantages of high labor intensity, low work efficiency, and inability to quickly cover large areas. In addition, the quality of manual visual inspection is unstable and easily affected by environmental factors. Under the influence of environmental factors such as obstructed vision, the inspection results are inaccurate and it is difficult to find hidden problems.

[0003] Compared with traditional manual visual inspection, the detection of the number of embedded sleeves based on the target detection algorithm can quickly process a large amount of image data, automatically identify and count the embedded sleeves, greatly improving the detection efficiency. Compared with traditional manual visual inspection, this method can save a lot of manpower and time costs; and in environments where the line of sight of manual visual inspection is obstructed, the successful identification of embedded parts can be ensured by arranging sensors, reducing the possibility of human errors and missed detections.

[0004] Due to the complexity of the prefabricated tunnel segment environment in this application, the existing manual visual inspection cannot meet the dual requirements of accuracy and speed. Therefore, a prefabricated tunnel segment embedded sleeve detection method based on a target detection algorithm is proposed. Summary of the invention

[0005] The purpose of the present invention is to overcome the shortcomings of the above technologies and provide a method for detecting embedded sleeves of prefabricated tunnel segments based on a target detection algorithm, which takes into account both high discrimination speed and high discrimination accuracy and is suitable for detecting the number of embedded sleeves.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] In a first aspect, the present invention provides a method for detecting the number of pre-embedded sleeves of a prefabricated tunnel segment based on a target detection algorithm, the detection method comprising the following steps:

[0008] Use a panoramic camera to obtain images of tunnel segment molds with and without pre-embedded sleeves installed before pouring; convert the size of the mold image into a standard size of 224*224 to obtain a standard size image;

[0009] Each pixel of the acquired standard format image is multiplied by the same pixel weight to obtain a new image to highlight the embedded sleeve, and the pixel weight is set to 1.25;

[0010] All new images constitute a data set, which is divided into a training set and a test set. The images in the training set that contain embedded sleeves are labeled, and the embedded sleeves are marked. Secondly, the data set is mean filtered, and the training set is enhanced.

[0011] Mean filtering: Use a 3*3 window to slide on the new image multiplied by the pixel weight, calculate the average value of the pixels in each window, and replace the pixel in the center of the window with this average value;

[0012] The process of data enhancement is as follows: the image after mean filtering is rotated clockwise by 30°, 60° and 90° respectively, and then the image is randomly erased, that is, a 50*50 rectangular frame filled with random values ​​is covered in the image to simulate the embedded sleeve partially blocked by steel, and finally the enhanced data set is formed;

[0013] A KAN-DETR target detection model is constructed. The KAN-DETR target detection model includes a backbone network ResNet-kan, a hybrid encoder, and a decoder with an auxiliary prediction head. The image is input into the backbone network ResNet-kan, and feature extraction is performed to obtain feature maps of multiple layers and different scales. The multi-scale features are then processed using a hybrid encoder. The output of the hybrid encoder is connected to the DIoU-aware query selection, and a fixed number of features are selected as the initial target query of the decoder. The decoder then generates a bounding box and a confidence score. The output of the decoder is converted into a probability distribution through a feedforward neural network, thereby obtaining a final prediction result, and then judging whether the embedded sleeve is missing.

[0014] The enhanced dataset is used to train the KAN-DETR target detection model for detecting the number of embedded sleeves in prefabricated tunnel segments.

[0015] Furthermore, in the backbone network ResNet-kan, the input image first passes through a convolution layer with a 7*7 convolution kernel, the number of output channels of the convolution layer is 64, and then passes through a 3*3 pooling layer. The data after pooling enters four convolution groups with different numbers of channels. These four convolution groups are composed of two 1*1 convolution kernels and one 3*3 convolution kernel. These convolution groups are repeated different times. The first convolution group is repeated 3 times, and the number of output channels is 256. The second convolution group is repeated 4 times, and the number of output channels is 256. The number of channels is 512, the third convolution group is repeated 6 times, and the number of output channels is 1024, and the fourth convolution group is repeated 3 times, and the number of output channels is 2048; the feature image G1(x) after convolution by the four convolution groups is used as the input of the MultKAN module, and then processed by the average pooling layer, the fully connected layer, and the activation function softmax to obtain output features with step sizes of 8, 16, and 32, respectively, which are denoted as S3, S4, and S5, respectively. S3, S4, and S5 are used as the input features of the hybrid encoder.

[0016] Furthermore, in the MultKAN module, the feature image G1(x) enters the first 1*1 convolution kernel with 512 convolution kernels, reducing the number of channels to 512, and then enters the second 3*3 convolution kernel with a step size of 2, further reducing the dimension of the feature image to obtain G2(x), and then enters the MultKAN layer. The output of the MultKAN layer MultKAN(G1(x)) is feature added with G2(x) to obtain the output of the MultKAN module.

[0017] Furthermore, the decoder includes a deformable attention mechanism, a normalization layer, a self-attention mechanism, a normalization layer, GR-kan, and a normalization layer connected in sequence; wherein the calculation formula of the deformable attention mechanism is:

[0018]

[0019] Among them, y is the output feature map, p0 is the point in the output space, and p n is the relative position of the convolution kernel; ω(p n ) is the weight of the convolution kernel; x is the input feature map; Δp n is the learned offset;

[0020] The calculation method of GR-kan is as follows: the internal univariate function Internal(x) maps the input real value x to an output value between 0 and 1 in the form of a sigmoid function.

[0021] The external single variable function External(x) connects a set of known data points through piecewise linear approximation to form a smooth curve. The calculation formula is:

[0022]

[0023] Where n1 represents the number of internal univariate functions Internal(x), c i are the coefficients optimized during training, P(x), Q(x) are polynomials; the output of the GR-Kan network is F(x), which is calculated as:

[0024] F(x)=ω(Internal(x)+External(x))

[0025] Where ω is the weight, Internal(x) is the internal univariate function, and spline(x) is the external univariate function.

[0026] Furthermore, the feedforward neural network includes linear combination, BN normalization processing and activation function ReLU, and finally obtains the number of pre-embedded sleeves.

[0027] The calculation formula of the ReLU activation function is:

[0028]

[0029] in, is the normalized data. When it is positive or zero, the ReLU function directly outputs The number of non-zero features is counted as the sleeve; and when When it is a negative number, the ReLU function outputs 0.

[0030] In a second aspect, the present invention provides a system for detecting the number of prefabricated tunnel segment embedded sleeves based on a target detection algorithm, the system comprising:

[0031] Image acquisition module, using a panoramic camera to obtain images of tunnel segment molds before pouring;

[0032] An image processing module is used to annotate and enhance the mold image of the image acquisition module;

[0033] The KAN-DETR target detection model is connected to the image processing module. The KAN-DETR target detection model is trained using the data output by the image processing module. The trained KAN-DETR target detection model is used to detect the embedded sleeve and output a prediction map.

[0034] Statistics module: used to count the number of non-zero features in the prediction image, and the number of non-zero features is the number of embedded sleeves detected in the current image;

[0035] The early warning and feedback module compares the results output by the statistical module with the preset number of pre-embedded sleeves on the tunnel segment. If the results output by the statistical module do not match the preset number, an alarm will be issued and the staff will be reminded to add or reduce the pre-embedded sleeves in time.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] In the present invention, the image of the tunnel segment mold is used as the input of target detection. The target detection uses the KAN-DETR target detection model. The decoder generates a bounding box and a confidence score. Finally, the output of the model is converted into a specific quantity to determine whether the embedded sleeve is missing.

[0038] The present invention not only improves the model prediction speed but also improves the model accuracy. In the embodiment, the recall rate is not less than 98%, the FPS is 78, and the accuracy is 98%, which meets the accuracy and real-time requirements when counting the number of embedded sleeves. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 The figure is a schematic diagram of the position of the embedded sleeve in the prefabricated tunnel segment according to an embodiment of the present invention, and the position indicated by the arrow in the figure is the installation position of the embedded sleeve.

[0040] Figure 2 It is a structural schematic diagram of the KAN-DETR target detection model in the present invention.

[0041] Figure 3 It is a structural diagram of the backbone network ResNet-kan in the present invention.

[0042] Figure 4 It is a flow chart of the MultKAN module in the present invention.

[0043] Figure 5 It is a schematic diagram of the structure of the MultKAN layer in the present invention.

[0044] Figure 6 It is a flow chart of the intra-scale feature interaction module AIFI of the present invention.

[0045] Figure 7 It is a flow chart of the mechanism of DIoU-aware query selection in the present invention.

[0046] Figure 8 It is a schematic diagram of the structure of the decoder in the present invention.

[0047] Fig. 9 It is a schematic diagram of the structure of the feedforward neural network in the present invention. DETAILED DESCRIPTION

[0048] In order to more clearly describe the technical problems, technical solutions and advantages of the present invention, the following will be described in detail with reference to the drawings and embodiments. It should be noted that these embodiments are only used to illustrate the principles and application scope of the present invention and should not be regarded as limiting the present invention.

[0049] Example 1

[0050] The method for detecting the number of prefabricated tunnel segment embedded sleeves based on the target detection algorithm in this embodiment includes the following steps:

[0051] (1) Using the image acquisition module, obtain the image of the tunnel segment mold before pouring;

[0052] (1.1) Using a panoramic camera, obtain images of the segment mold with and without pre-embedded sleeves installed. The images taken by the panoramic camera can reflect the pre-embedded sleeves on each surface of the tunnel segment. By counting the number of pre-embedded sleeves in a single image, it can be determined whether the number of pre-embedded sleeves on the tunnel segment is qualified.

[0053] (1.2) Convert the mold image format into a standard format of 224*224 to obtain a standard format image;

[0054] (1.3) Highlighting the embedded sleeve: Multiply each pixel of the acquired standard format image by the same pixel weight to obtain a new image. Here, the pixel weight is set to 1.25.

[0055] The detection of the number of embedded sleeves in prefabricated tunnel segments is completed in the prefabrication factory. The detection environment is relatively fixed. The embedded sleeves are generally silver-white and steel is gray. Under the same brightness, the silver-white embedded sleeves are brighter than the gray steel. In order to highlight the embedded sleeves, the pixel weight is set to 1.25. When the pixel weight is greater than 1, the contrast is enhanced to make the black and white of the image more obvious.

[0056] (1.4) All new images constitute a data set, which includes images with pre-embedded sleeves and images without pre-embedded sleeves. The two types of images are allocated as training sets and test sets in a ratio of 7:3, and the images with pre-embedded sleeves in the training set are labeled to mark the pre-embedded sleeves. Secondly, the data set is mean filtered, and the training set is data enhanced.

[0057] Mean filtering: Use a 3*3 window to slide on the new image multiplied by the pixel weight, calculate the average value of the pixels in each window, and replace the pixels in the center of the window with this average value; reduce random noise in the image;

[0058] The process of data enhancement is as follows: the image after mean filtering is rotated clockwise by 30°, 60° and 90° respectively, and then the image is randomly erased, that is, a 50*50 rectangular frame filled with random values ​​is covered in the image to simulate the embedded sleeve part blocked by steel, and finally an enhanced data set is formed. The number of enhanced data sets is 1500.

[0059] Constructing a KAN-DETR target detection model, wherein the KAN-DETR target detection model includes a backbone network ResNet-kan, a hybrid encoder, and a decoder with an auxiliary prediction head;

[0060] The mold image is used as the input of the backbone network ResNet-kan. It first passes through the convolution layer with a 7*7 convolution kernel. The number of output channels of the convolution layer is 64, and then passes through the 3*3 pooling layer. After pooling, the data enters four convolution groups. These four convolution groups are composed of two 1*1 convolution kernels and one 3*3 convolution kernel. These convolution groups are repeated different times. The first convolution group is repeated 3 times, and the number of output channels is 256. The second convolution group is repeated 4 times, and the number of output channels is 51. 2. The third convolution group is repeated 6 times, and the number of output channels is 1024. The fourth convolution group is repeated 3 times, and the number of output channels is 2048. The feature image G1(x) after convolution by the four convolution groups is used as the input of the MultKAN module, and then processed by the average pooling layer, the fully connected layer, and the activation function softmax to obtain output features with step sizes of 8, 16, and 32, respectively, which are denoted as S3, S4, and S5, respectively. S3, S4, and S5 are used as the input features of the hybrid encoder.

[0061] In the MultKAN module, the feature image G1(x) enters the first convolution kernel, the convolution kernel size is 1*1, the step size is 1; the number of convolution kernels is 512, the number of channels is reduced to 512, and then enters the second 3*3 convolution kernel, the step size is 2, and the number of convolution kernels is 512; the feature image is further reduced in dimension to obtain G2(x), and then enters the MultKAN layer,

[0062] In the MultKAN layer, the residual activation function is first used. The nonlinear activation is calculated as follows:

[0063]

[0064] Among them, B(x) is the B-spline function, n is the number of B-spline functions, and the output obtained by the residual activation function is After addition and multiplication, the residual activation function is input to the next layer. , the activation functions of different layers are calculated as follows:

[0065]

[0066] Where m represents the total number of MultKAN layers, and the symbol ° represents the composite operation between functions, that is, the output of one function becomes the input of the next function.

[0067] The output of the MultKAN layer, MultKAN(G1(x)), is added to G2(x) to further improve the generalization ability of the ResNet-kan network.

[0068] The hybrid encoder includes an intra-scale feature interaction module AIFI and a multi-scale feature fusion module CCFM. The input feature S5 is used as the input of the intra-scale feature interaction module AIFI (encoder). S5 first obtains a reconstructed multi-dimensional feature image through a multi-head self-attention mechanism. The specific transformation formula is as follows:

[0069] Q=K=V=flatten(S5)

[0070] F5=Reshape(Attn(Q,K,V))

[0071] Among them, flatten is the process of converting the multi-dimensional array of S5 image features into a one-dimensional array, that is, merging each row or column of the S5 image features into a separate sequence, which is the flattening feature operation; Attn refers to processing the vectors Q, K and V with a multi-head self-attention mechanism, and Reshape refers to the deformation operation, which is to reshape the result of the multi-head self-attention mechanism into a multi-dimensional image feature F5, that is, to reconstruct the feature map.

[0072] S3, S4, and F5 are used as the input of the multi-scale feature fusion module CCFM to perform multi-scale feature fusion. The specific process is (see Figure 2 ):

[0073] After 1*1 convolution, BN normalization and activation function calculation, F5 is used together with S4 as the input of fusion block 1. After fusion, the first fusion result of fusion block 1 is obtained. The mechanism of the fusion block is as follows:

[0074] Features of different scales are used as the input of the fusion block and first convolved with a 1*1 convolution kernel. Then, the features are concatenated by element-wise addition to obtain a new feature map as the output of the fusion block.

[0075] The first fusion result is processed by 1*1 convolution, BN normalization and activation function, and then used together with S3 as the input of fusion block 2 to obtain the output of fusion block 2, which is recorded as the second fusion result. The output of fusion block 2 is processed again by image processing (3*3 convolution, BN normalization and activation function) and then used as the input of fusion block 1 again. After fusion, the second output of fusion block 1 is obtained. The second output of fusion block 1 is processed by image processing (3*3 convolution, BN normalization and activation function) and F5 is processed by image processing (1*1 convolution, BN normalization and activation function) as the input of fusion block 3. After fusion, the output of fusion block 3 is obtained, which is recorded as the third fusion result. Finally, the second output of fusion block 1, the output of fusion block 2 and the output of fusion block 3 are concatenated by element addition to obtain a new feature map, and the output output of the hybrid encoder is obtained. The formula is as follows:

[0076] Output = CCFM ({S3, S4, F5})

[0077] A fixed number of image features are selected from the hybrid encoder output sequence as the initial object query for the decoder. DIoU-aware query selection is achieved by constraining the model to produce high classification scores for features with high DIoU scores and low classification scores for features with low DIoU scores during training. DIOU is calculated as follows:

[0078]

[0079] Where IOU is the overlap between the predicted bounding box and the true bounding box, c represents the center point distance, and d represents the diagonal distance. The prediction box corresponding to the first K hybrid encoder features (K is 300 in this embodiment) selected by the model based on the classification score has a higher classification score and a higher DIoU score. The optimization objectives are as follows:

[0080]

[0081] in and y represent the prediction and true value respectively, c and b represent the category and bounding box respectively, and Represent the predicted category and predicted bounding box respectively; and y represent the prediction and true value respectively, c and b represent the category and bounding box respectively, and Represent the predicted category and predicted bounding box respectively; L box , L cls , L are the bounding box loss, category loss and total loss respectively.

[0082] Finally, the boxes with high DIoU scores and high classification scores enter the decoder, and pass through the deformable attention mechanism, normalization layer, self-attention mechanism, normalization layer, GR-kan, and normalization layer in sequence to obtain the output of the decoder. The calculation formula of the deformable attention mechanism is as follows:

[0083]

[0084] Among them, y is the output feature map, p0 is the point in the output space, and p n is the relative position of the convolution kernel. ω(p n ) is the weight of the convolution kernel. x is the input feature map, Δp n is the learned offset.

[0085] Among them, the calculation formula of the self-attention mechanism is as follows:

[0086]

[0087] Q, K, and V are query, key, and value vectors respectively, d k is the dimension of the key vector, used to scale the dot product for improved numerical stability.

[0088] The calculation method of GR-kan is as follows: the internal univariate function BNternal(x) maps the input real value x to an output value between 0 and 1 in the form of a sigmoid function, and the calculation method is as follows:

[0089] Internal(x)=x / (1+e -x )

[0090] The external univariate function External(x) connects a set of known data points through piecewise linear approximation to form a smooth curve. The calculation method is as follows:

[0091]

[0092] Where n1 represents the number of internal univariate functions BNternal(x), c i are the coefficients optimized during training, P(x), Q(x) are polynomials. The output of the GR-Kan network is F(x), which is calculated as follows:

[0093] F(x)=ω(Internal(x)+External(x))

[0094] Where ω is the weight, the internal univariate function is BNternal(x), and the external univariate function is splBNe(x).

[0095] The output of the decoder is used as the input of the feedforward neural network. After linear combination, normalization and activation function ReLU processing in the feedforward neural network, a prediction map is obtained. In the prediction map, 0 indicates a non-embedded sleeve, and non-0 indicates an embedded sleeve. Finally, the number of non-0 in the prediction map is counted to obtain the number of embedded sleeves. The linear combination is: the input is scaled and summed by weight, and then offset by the deviation. The specific formula is:

[0096] z=w1x1+w2x2+....+w n x n +b

[0097] Where w is the weight of data x, b is the bias; z is the linear combination result; n represents the number of elements in data x (a feature of the decoder output);

[0098] The specific process of BN normalization is:

[0099] 1) Calculate the batch mean μ B and variance σ 2 B

[0100]

[0101] Where m is the number of samples, x i is the sample data;

[0102] 2) Normalization

[0103]

[0104] in, is the normalized data, ∈ is a small positive constant used to prevent numerical instability caused by the denominator being zero;

[0105] The calculation formula of the ReLU activation function is:

[0106]

[0107] Among them, when When it is positive or zero, the ReLU function directly outputs The number of non-zero features is counted as the sleeve; and when When it is a negative number, the ReLU function outputs 0.

[0108] Example 2

[0109] This embodiment is a detection method for embedded sleeves of prefabricated tunnel segments based on the target detection algorithm (KAN-DETR). The KAN-DETR neural network consists of a backbone network ResNet-kan, a hybrid encoder, and a decoder with an auxiliary prediction head. The image is input into the backbone network ResNet-kan, and the output characteristics with step sizes of 8, 16, and 32 are obtained as the input of the hybrid encoder, and S3, S4, and S5 are used to mark them respectively. S5 is used as the input of the AIFI module, and F5 is obtained through intra-scale interaction. S3, S4, and F5 are used as the input of the CCFM module for multi-scale feature fusion. The processing mechanism of image features is as follows:

[0110] The image features are first convolved with the convolution kernel to further obtain new image features, then normalized by the normalization technology BN, and then calculated by the activation function ReLU to obtain higher-level abstract features. The image feature F5 is convolved with 1*1 convolution, normalized, and activated together with S4 as the input of fusion block 1. After fusion, the first output of fusion block 1 (the first fusion result) is obtained. The mechanism of the fusion block is as follows: the features of different scales are first convolved with 1*1 convolution kernel as the input of the fusion block, and then fused by element addition to obtain a new feature map as the output of the fusion block.

[0111] The first fusion result is processed by 1*1 convolution, BN normalization and activation function, and then used together with S3 as the input of fusion block 2 to obtain the output of fusion block 2, which is recorded as the second fusion result. The output of fusion block 2 is processed again by image processing (3*3 convolution, BN normalization and activation function) and then used as the input of fusion block 1 again. After fusion, the second output of fusion block 1 is obtained. The second output of fusion block 1 is processed by image processing (3*3 convolution, BN normalization and activation function) and the result of image processing (1*1 convolution, BN normalization and activation function) of F5 as the input of fusion block 3. After fusion, the output of fusion block 3 is obtained, which is recorded as the third fusion result. Finally, the second output of fusion block 1, the output of fusion block 2 and the output of fusion block 3 are concatenated by element addition to obtain a new feature map.

[0112] A fixed number of image features are selected from the hybrid encoder output sequence as the initial object query for the decoder. DIoU-aware query selection is achieved by constraining the model to produce high classification scores for features with high DIoU scores and low classification scores for features with low DIoU scores during training. Therefore, the prediction boxes corresponding to the first K hybrid encoder features selected by the model based on the classification scores have high classification scores and high DIoU scores. After optimization by the constrained model, the prediction boxes with high DIoU scores and high classification scores enter the decoder, where they pass through the deformable attention mechanism, normalization layer, self-attention mechanism, normalization layer, GR-kan, and normalization layer in sequence to obtain the output of the decoder, the precise bounding box and its corresponding confidence score. The output of the decoder is used as the input of the feedforward neural network. The data is first linearly combined in the feedforward neural network, and the normalized data is then nonlinearly activated. The output of the model is converted into quantity through the activation function ReLU processing, thereby obtaining the final detection result.

[0113] The hybrid encoder consists of an intra-scale feature interaction module AIFI and a multi-scale feature fusion module CCFM. Intra-scale interaction is performed in the AIFI module, and multi-scale feature fusion is performed in the CCFM module.

[0114] In the decoder, the predicted box after DIOU perception query is input into the decoder, and passes through the deformable attention mechanism, normalization layer, self-attention mechanism, normalization layer, GR-kan, normalization layer in sequence to obtain the output of the decoder, the precise bounding box and its corresponding confidence score.

[0115] Compared with the conventional manual visual inspection of the number of embedded sleeves, which has the disadvantages of high labor intensity, low work efficiency, unstable quality, and inaccurate inspection results due to environmental factors such as obstructed vision, the prefabricated tunnel segment embedded sleeve number detection method based on KAN-DETR of the present invention greatly reduces labor costs, improves the detection rate of the embedded sleeve number, and ensures the detection quality, thereby realizing high-precision quantity detection.

[0116] The labeled training set is read into the KAN-DETR target detection model, and the target detection model is trained. All images are iterated and trained in turn. After 40,000 iterations, the training model reaches convergence, that is, when the model training gradient is close to 0 (less than 0.01 can be considered close to 0), the training is stopped, and the optimal network parameters are extracted for prediction. If the training gradient is not close to 0, the weight parameters are adjusted. Model calibration stage: 1) Input the test set data into the KAN-DETR target detection model to obtain the detection results; 2) Compare the actual parameters of the test target with the preliminary detection results to obtain the calibrated KAN-DETR model; When predicting, the model first loads the trained parameters and loads the input image from the test set. The prediction map obtained by the trained KAN-DETR target detection model is used to detect the number of embedded sleeves.

[0117] The training set is used to train the model, and the test set is used to evaluate the model performance under different parameter settings, so as to select the best parameter settings and make the model have better generalization ability. At the same time, the data set is enhanced by rotating and randomly erasing the pictures, which is conducive to identifying the embedded sleeves at different angles and improving the overall performance of the model.

[0118] Combining the MultKAN module with the backbone network enables the backbone network to better capture the complex relationships in the data and enhances the interpretability of the backbone network. Using GR-kan in the decoder improves the accuracy and speed of the overall model.

[0119] The performance of different models was tested using the enhanced data set established in Example 1. The test results are shown in the following table:

[0120]

[0121] Among them, Recall is the recall rate, which is the ratio of positive samples correctly identified by the model to all samples that are actually positive, and is a measure of the model's ability to identify positive samples. Precision is the accuracy rate, which is the ratio of samples predicted by the model to samples that are actually positive, and is a measure of the accuracy of the model's prediction results. FPS indicates how many frames of images the model can process in one second. Compared with the existing common target detection algorithms, the present invention not only improves the model prediction speed but also improves the model accuracy. The recall rate of the present invention is not less than 98%, the FPS is 78, and the accuracy is 98%, which meets the accuracy and real-time requirements.

[0122] Example 3

[0123] This embodiment is based on the KAN-DETR prefabricated tunnel segment embedded sleeve quantity detection system, including:

[0124] Image acquisition module, using a panoramic camera to obtain images of tunnel segment molds before pouring;

[0125] An image processing module is used to annotate and enhance the mold image of the image acquisition module;

[0126] The KAN-DETR target detection model is connected to the image processing module. The KAN-DETR target detection model is trained using the data output by the image processing module. The trained KAN-DETR target detection model is used to detect the embedded sleeve and output a prediction map.

[0127] Statistics module: used to count the number of non-zero features in the prediction image, and the number of non-zero features is the number of embedded sleeves detected in the current image;

[0128] The early warning and feedback module compares the results output by the statistical module with the preset number of pre-embedded sleeves on the tunnel segment. If the results output by the statistical module do not match the preset number, an alarm will be issued and the staff will be reminded to add or reduce the pre-embedded sleeves in time.

[0129] The training process of the KAN-DETR target detection model is: start with random initial KAN-DETR network data, load the training set image, input it into the KAN-DETR target detection model, predict the image and calculate the loss function error; the training gradient is close to 0, and the trained KAN-DETR target detection model network parameters are obtained. If the gradient is not close to 0, the error back propagation is performed to adjust the model parameters, and the loaded training set image is returned. The test set is input into the trained KAN-DETR target detection model. The trained KAN-DETR target detection model can be used to realize the detection of the number of embedded sleeves; the trained KAN-DETR target detection model is used to obtain the prediction map, and then the number of embedded sleeves in the mold image is detected.

[0130] In actual use, the panoramic camera captures the image of the tunnel segment mold before pouring. After the mold image before pouring is processed by standard size, pixel weight and mean filtering, it is input into the trained KAN-DETR target detection model and the corresponding prediction map is output;

[0131] The number of non-zero features in the statistical prediction graph is compared with the number of embedded sleeves on the preset tunnel segment. If the number of embedded sleeves is identified to be correct, the correct number will be displayed; if the number of embedded sleeves is identified to be insufficient or excessive, an alarm will be issued to remind the staff to add or reduce the embedded sleeves in time.

[0132] Example 4

[0133] The hardware devices used in the detection system of this embodiment include the following components:

[0134] Image sensor: The image sensor is used to capture images, convert optical images into electrical signals, and then form a digital image. The image sensor uses a panoramic camera to meet the needs of the scene.

[0135] Processor: As the core component of the present invention, the processor is responsible for controlling and managing the operation of the entire system, including data acquisition, data processing, image recognition and other functions, and needs to have sufficient computing power and parallel processing capabilities to meet real-time requirements. The processor can be in different forms such as single-chip microcomputers, microprocessors, computers, etc. to meet the needs of different application scenarios;

[0136] Memory: Memory can be used to store collected data and historical data for subsequent processing and analysis. It has the characteristics of high speed, high reliability and scalability to meet the needs of long-term stable operation of the system;

[0137] Database: Use database to store and manage collected data, historical data, analysis results and other information;

[0138] Network interface: used for data exchange and communication, with the characteristics of high speed, high stability and high security, ensuring the reliability and security of data transmission.

[0139] The processor is configured to execute computer executable instructions. When the computer executable instructions are executed by the processor, the various steps of the method for detecting the number of embedded sleeves of prefabricated tunnel segments based on KAN-DETR are implemented.

[0140] A computer program is stored in the memory, and the computer program can be executed by the processor to implement each step of the method for detecting the number of embedded sleeves of prefabricated tunnel segments based on KAN-DETR.

[0141] The database is configured to store and manage data of a computer application, including various data types and structures, which are applied to various steps of a method for detecting the number of embedded sleeves of prefabricated tunnel segments based on KAN-DETR.

[0142] The network interface realizes communication and data transmission between computers. The network interface can provide various communication protocols and data transmission methods to meet the communication and data transmission requirements of different application scenarios and different needs, and is applied to each step of the prefabricated tunnel segment embedded sleeve quantity detection method based on KAN-DETR.

[0143] The present invention is mainly used for detecting the number of embedded sleeves in a prefabricated tunnel segment before casting, and uses a panoramic camera to automatically identify the number of embedded sleeves.

[0144] The present invention aims to solve the problem that the current manual detection of the number of embedded sleeves is slow and affects the construction progress, and automatically identifies the number of embedded sleeves through the KAN-DETR target detection model. For the number of embedded sleeves, the target detection model can quickly identify and automatically alarm, and promptly remind the staff to complete the increase or decrease of the embedded sleeves. Compared with the existing methods, this technical solution has the following advantages and application prospects: improving the efficiency of detecting the number of embedded sleeves, and providing the possibility of pursuing more intelligent prefabricated tunnel segments. This is of great significance for prefabricated tunnel segments and has broad application prospects.

[0145] In the description of this specification, the described specific features, structures, or characteristics may be combined in an appropriate manner in any one or more embodiments or examples.

[0146] It should be understood that each part of the present invention can be implemented by hardware, software or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or hardware stored in a memory and executed by a suitable instruction execution system.

[0147] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0148] Any matters not described in the present invention are applicable to the prior art.

Claims

1. A method for detecting the number of pre-embedded sleeves of prefabricated tunnel segments based on a target detection algorithm, characterized in that: The detection method comprises the following steps: Use a panoramic camera to obtain images of tunnel segment molds with and without pre-embedded sleeves installed before pouring; convert the size of the mold image into a standard size of 224*224 to obtain a standard size image; Each pixel of the acquired standard format image is multiplied by the same pixel weight to obtain a new image to highlight the embedded sleeve, and the pixel weight is set to 1.25; All new images constitute a data set, which is divided into a training set and a test set. The images in the training set that contain embedded sleeves are labeled, and the embedded sleeves are marked. Secondly, the data set is mean filtered, and the training set is enhanced. Mean filtering: Use a 3*3 window to slide on the new image multiplied by the pixel weight, calculate the average value of the pixels in each window, and replace the pixel in the center of the window with this average value; The process of data enhancement is as follows: the image after mean filtering is rotated clockwise by 30°, 60° and 90° respectively, and then the image is randomly erased, that is, a 50*50 rectangular frame filled with random values ​​is covered in the image to simulate the embedded sleeve partially blocked by steel, and finally the enhanced data set is formed; A KAN-DETR target detection model is constructed. The KAN-DETR target detection model includes a backbone network ResNet-kan, a hybrid encoder, and a decoder with an auxiliary prediction head. The image is input into the backbone network ResNet-kan, and feature extraction is performed to obtain feature maps of multiple layers and different scales. The multi-scale features are then processed using a hybrid encoder. The output of the hybrid encoder is connected to the DIoU-aware query selection, and a fixed number of features are selected as the initial target query of the decoder. The decoder then generates a bounding box and a confidence score. The output of the decoder is converted into a probability distribution through a feedforward neural network, thereby obtaining a final prediction result, and then judging whether the embedded sleeve is missing. The enhanced dataset is used to train the KAN-DETR target detection model for detecting the number of embedded sleeves in prefabricated tunnel segments.

2. The detection method according to claim 1, characterized in that: In the backbone network ResNet-kan, the input image first passes through a convolution layer with a 7*7 convolution kernel, and the number of output channels of the convolution layer is 64, and then passes through a 3*3 pooling layer. After pooling, the data enters four convolution groups with different numbers of channels. These four convolution groups are composed of two 1*1 convolution kernels and one 3*3 convolution kernel. These convolution groups are repeated different times. The first convolution group is repeated 3 times, and the number of output channels is 256. The second convolution group is repeated 4 times, and the number of output channels is 512, the third convolution group is repeated 6 times, the number of output channels is 1024, the fourth convolution group is repeated 3 times, the number of output channels is 2048; the feature image G1(x) after convolution by the four convolution groups is used as the input of the MultKAN module, and then processed by the average pooling layer, the fully connected layer, and the activation function softmax to obtain output features with step sizes of 8, 16, and 32, respectively, which are denoted as S3, S4, and S5, respectively. S3, S4, and S5 are used as the input features of the hybrid encoder.

3. The detection method according to claim 2, characterized in that: In the MultKAN module, the feature image G1(x) enters the first 1*1 convolution kernel with 512 convolution kernels, reducing the number of channels to 512, and then enters the second 3*3 convolution kernel with a step size of 2, further reducing the dimension of the feature image to obtain G2(x), and then enters the MultKAN layer. The output of the MultKAN layer, MultKAN(G1(x)), is feature added with G2(x) to obtain the output of the MultKAN module.

4. The detection method according to claim 1, characterized in that: The decoder includes a deformable attention mechanism, a normalization layer, a self-attention mechanism, a normalization layer, GR-kan, and a normalization layer connected in sequence; wherein the calculation formula of the deformable attention mechanism is: Among them, y is the output feature map, p0 is the point in the output space, and p n is the relative position of the convolution kernel; ω(p n ) is the weight of the convolution kernel; x is the input feature map; Δp n is the learned offset; The calculation method of GR-kan is as follows: the internal univariate function Internal(x) maps the input real value x to an output value between 0 and 1 in the form of a sigmoid function. The external single variable function External(x) connects a set of known data points through piecewise linear approximation to form a smooth curve. The calculation formula is: Where n1 represents the number of internal univariate functions Internal(x), c i are the coefficients optimized during training, P(x), Q(x) are polynomials; the output of the GR-Kan network is F(x), which is calculated as: F(x)=ω(Internal(x)+External(x)) Where ω is the weight, Internal(x) is the internal univariate function, and spline(x) is the external univariate function.

5. The detection method according to claim 1, characterized in that: The feedforward neural network includes linear combination, BN normalization processing and activation function ReLU, and finally obtains the number of pre-embedded sleeves. The calculation formula of the ReLU activation function is: in, is the normalized data. When it is positive or zero, the ReLU function directly outputs The number of non-zero features is counted as the sleeve; and when When it is a negative number, the ReLU function outputs 0.

6. A system for detecting the number of prefabricated tunnel segment embedded sleeves based on a target detection algorithm, characterized in that: The system comprises: Image acquisition module, using a panoramic camera to obtain images of tunnel segment molds before pouring; An image processing module is used to annotate and enhance the mold image of the image acquisition module; The KAN-DETR target detection model is connected to the image processing module. The KAN-DETR target detection model is trained using the data output by the image processing module. The trained KAN-DETR target detection model is used to detect the embedded sleeve and output a prediction map. Statistics module: used to count the number of non-zero features in the prediction image, and the number of non-zero features is the number of embedded sleeves detected in the current image; The early warning and feedback module compares the results output by the statistical module with the preset number of pre-embedded sleeves on the tunnel segment. If the results output by the statistical module do not match the preset number, an alarm will be issued and the staff will be reminded to add or reduce the pre-embedded sleeves in time.

Citation Information

Patent Citations

  • Underground cable pipeline scene foreign matter identification and classification method

    CN116311006A

  • Express item security check method

    CN117218081A

  • Target detection method based on OTS-DETR lightweight model

    CN118262087A

  • Lightweight road defect detection method based on dynamic deformable attention mechanism

    CN118279562A

  • Method for detecting internal cavity and non-compact defect of tunnel lining based on GPR (General Purpose Register) data

    CN118628452A