Lightweight intelligent detection method for railway tunnel lining diseases
By constructing a lightweight convolutional neural network for railway tunnel lining disease detection, the problems of low efficiency and high cost in the existing technology are solved, and efficient and accurate detection of multiple diseases is achieved.
Patent Information
- Application Number
- CN202411902217.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-23
AI Technical Summary
The existing railway tunnel lining disease detection methods rely on manual identification, with low efficiency and inconsistent standards. Machine learning methods consume high calculations and low accuracy in long tunnels. The existing deep learning models have large parameters and high costs, making it difficult to achieve lightweight detection.
A convolutional neural network is used to build a lightweight model, and the image is processed through adaptive scaling and filling, combining efficient feature extraction, lightweight feature fusion and feature enhancement modules to achieve end-to-end detection of lining diseases, including the identification of evacuated structures and steel bars.
It improves the intelligence and efficiency of railway tunnel lining disease detection, reduces hardware requirements, and realizes high-precision detection of multiple types of diseases, which is suitable for long-distance tunnel detection.
Smart Images

Figure CN120014431A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of image detection, and in particular to a lightweight intelligent detection method for railway tunnel lining defects. Background Art
[0002] The stability of railway tunnels is related to the safety of railway operations and the safety of people's lives and property. Therefore, it is of great significance to conduct quality inspections on railway tunnels. Among them, tunnel lining is an important inspection item of quality inspection engineering, and its disease detection technology research and application are particularly critical.
[0003] At present, based on the non-destructive detection properties of tunnel lining radar, tunnel lining disease detection technology is usually carried out based on geological radar images, mainly including subjective identification methods based on manual search, machine learning methods based on feature extraction engineering, and intelligent detection methods based on deep learning. The subjective identification method mainly conducts subjective searches for areas that meet typical disease characteristics based on the radar imaging mechanism. For narrow and long tunnels, the workload of subjective identification is huge; and this method relies heavily on the personal experience of the identification personnel, resulting in inconsistent identification standards for various defects.
[0004] Machine learning methods extract image features containing texture, color, and edge information from the defective area in the radar image, train multiple classifiers for various defects, and then assist in manual identification. This type of method is highly dependent on feature extraction of the defective area. In long tunnel lining radar images, defective targets are often sparsely distributed. Therefore, if feature extraction and calculation are performed directly on a global scale, hardware and time consumption will increase dramatically. In addition, due to the influence of complex background interference in lining radar images, this type of method is prone to produce more false alarms, which is not conducive to accurate and rapid disease identification.
[0005] In recent years, with the rapid development of computer technology and artificial intelligence, deep learning-based methods have gradually been introduced into tunnel lining radar image analysis. The core step of this type of method is to detect the defect area with specific image features from the radar image. However, there are few studies directly used for lining disease detection, especially the research on network design based on typical disease samples and features, and the existing methods only detect a single type of disease, such as hollow disease identification. In addition, from the application perspective, the existing lining disease detection model has a huge number of parameters and occupies a large amount of memory. In engineering, it relies on the support of graphics processing units (GPUs) and parallel computing, which greatly increases the computer load capacity and engineering economic costs, and may reduce the stability of the detection equipment. The study of lightweight models is the key difficulty of current engineering applications. However, there is currently no lightweight model dedicated to railway tunnel lining disease detection.
[0006] In summary, the existing tunnel lining inspection projects are gradually in the transition stage from manual identification to intelligent detection. Researching high-performance and lightweight railway tunnel lining multi-type disease detection methods is of great significance for replacing manual identification, improving the accuracy of multi-type lining disease detection, and realizing low-energy consumption and low-cost engineering deployment and application. Summary of the invention
[0007] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes a lightweight intelligent detection method for railway tunnel lining defects, which is mainly used for railway tunnel lining defect detection, has a comparable conclusion to the manual identification result, can effectively replace manual intelligent identification, and can improve the efficiency and accuracy of long-distance tunnel lining quality detection.
[0008] In order to solve the above problems, the present invention provides a lightweight intelligent detection method for railway tunnel lining defects, comprising the following steps:
[0009] Step 1: Acquire multiple radar images of railway tunnel linings of a specific length, which contain multiple types of reflective structures such as secondary lining, steel bars, and primary supports;
[0010] Step 2: According to the mileage information of each image input data, the lining radar image is labeled, segmented, and the relative coordinates of the segmented image targets are transformed to construct an image dataset for training;
[0011] Step 3: For each segmentation map F i Scale normalization is performed, and adaptive scaling and padding are used to obtain a fixed input scale of c·c, where c = 640; assuming that the segmentation map F i The size of is a×b (a≥b), and the adaptive scaling factor is defined as the inverse of the ratio of the long side a to the fixed length c, that is, After scaling, the long side is scaled to c, and the short side length is λ·b. To convert to a normalized scale, the short sides on both sides are filled with 0, and the calculation formula for the single-side short side filling pixel size m is:
[0012]
[0013] Step 4: Construct a convolutional neural network for hollow structure and steel bar identification, including an efficient feature extraction network, a lightweight feature fusion network, a feature enhancement module, and a feature decoding structure; input the location and category labels of all segmentation map targets to train the neural network to obtain weight parameters;
[0014] Step 5: Perform the preprocessing operations of steps 2 and 3 on the image to be inspected, and then based on the training weight parameters of step 4 and the network structure inference, detect the categories and location parameters of the steel bars and the hollow structures in all the segmented images;
[0015] Step 6: According to the segmentation map F i The test results can be used to determine whether there is a lack of tendons;
[0016] Step 7: Based on the segmentation parameters of step 2, all segmentation map detection results are spliced to obtain a long detection result of the same size as the input data.
[0017] Preferably, step 2 includes the following sub-steps:
[0018] Step 2-1: The segmentation method adopts sliding window segmentation, and the step size s of the sliding window is set to 1000 pixels, the corresponding segmentation image mileage is about 10m, and the segmentation margin g is 200 pixels;
[0019] Step 2-2: Perform segmentation and shifting processing on the segmentation edge to ensure the integrity of the target and prevent the segmentation operation from damaging the steel bars and the empty targets; determine whether there is a target at each window edge, if the segmentation line passes through the target, perform shifting processing, if the segmentation line does not pass through the target, the segmentation area width remains unchanged; in this way, obtain several segmentation maps;
[0020] Step 2-3: To facilitate training, convert the initial marked target box position into the target box position coordinates in the segmentation map; assuming that the i-th segmentation map F i The original parameters of the target box in are (x1, y1, x2, y2) and the delayed pixels are e i , then the relative position (x1', y1', x2', y2') is calculated as follows:
[0021]
[0022] y′1=y1
[0023]
[0024] y′2=y2。
[0025] Preferably, step 5 includes the following sub-steps:
[0026] Step 5-1: Input the preprocessed image into the efficient feature extraction network to obtain the initial feature maps F at three scales i,80×80×128 、F i,40×40×256 、F i,20×20×512 ;
[0027] Step 5-2: Use a lightweight feature fusion network to perform cross-layer weighted fusion of the initial features to obtain a fusion feature map P of three scales i,80×80×64 , P i,40×40×128 , P i,20×20×256 ;
[0028] Step 5-3: Fusion feature map Pi,80×80×64 , P i,40×40×128 , P i,20×20×256 Implement context feature enhancement to obtain enhanced feature map
[0029] Step 5-4: Enhance the feature map Input feature decoding structure to obtain prediction vector; finally, the prediction vector contains (t x ,t y ,t w ,t h ), Sorce representing the confidence score of the current prediction box confi And Probability representing the probability of the target category class ; To normalize the output of the predicted target box, the coordinates of the center point of the predicted box (b x ,b y ,b w ,b h ) is calculated from the relative predicted coordinates of the corresponding grid (t x ,t y ,t w ,t h ) indicates that the calculation formula is as follows:
[0030] b x =c x +2σ(t x )-0.5
[0031] b y =c y +2σ(t y )-0.5
[0032] b w =p w ·(2σ(t w )) 2
[0033] b h =p h ·(2σ(t h )) 2 ;
[0034] Where (c x ,c y ) is the relative coordinate of the upper left corner of the grid responsible for predicting the target.
[0035] Preferably, step 5-1 includes the following sub-steps:
[0036] Step 5-1-1: Use depth-separable convolution, batch normalization, and SiLU function activation processing on the input image, and use the DW-CBS module to perform shallow feature extraction;
[0037] Step 5-1-2: Multiple stacking of GhostELA-DWC3 module and DW-CBS module for multi-scale feature extraction, GhostELA-DWC3 is an improved deep feature extraction module; the stacking method is: stacking DW-CBS module and GhostELA-DWC3 module twice to obtain feature map F i,80×80×128 , superimposed three times to obtain the feature map F i,40×40×256 , stack four times to get F i,20×20×512 .
[0038] Preferably, step 5-2 includes the following sub-steps:
[0039] Step 5-2-1: F i,20×20×512 The SPPF block is used for spatial pyramid pooling, and then GhostELA-DWC3, DW-CBS modules and nearest neighbor interpolation are used to obtain the double upsampled feature F. i,40×40×128 , this feature is consistent with F i,40×40×256 Splice in the channel direction to obtain the upsampled fusion feature map F i,40×40×384 ; In this way, another scale of upsampled fusion feature map F is obtained i,80×80×192 ;
[0040] Step 5-2-2: Feature map F i,80×80×192 The GhostELA-DWC3 module is further used to obtain the fusion feature map P i,80×80×64 ; Based on the idea of cross-layer weighted fusion, the channel weight W of the multi-layer feature map is calculated by fast normalization:
[0041]
[0042] Where w i is the learnable weight, ε = 0.0001 is used for stable propagation of gradients; at each w i Then apply SiLU activation function to ensure the weighted value w i >0; then, each normalized weight value is limited to between 0 and 1;
[0043] Step 5-2-3: Initial feature map, upsampled fusion feature map F i,80×80×128 And the fusion feature map P i,80×80×64 Perform weighted splicing in the channel direction to obtain the fusion feature map P i,40×40×128 ; In this way, the fusion feature map P is obtained i,20×20×256 .
[0044] Preferably, step 5-3 includes the following sub-steps:
[0045] Step 5-3-1: Fusion feature map P i,80×80×64Transformed into query Q, key K and value V vectors, the formula is as follows:
[0046] Q=P i,80×80×64 ·M q
[0047]
[0048] V=P i,80×80×64 ·M v ;
[0049] Among them, M q , M v are the embedding matrices used to transform query, key, and value vectors, respectively;
[0050] Step 5-3-2: Assume the center key of the context area is X cen , the size of the surrounding area is ×, where = 3, then × times of convolution can be calculated to obtain the key vector information of each surrounding area; the learned context key K Static Reflects the static information of the center and surroundings;
[0051] Step 5-3-3: Concatenate the context key and query vector to form a synthetic key [K Static ,]; self-attention encoding is performed by using two consecutive 1×1 convolutions:
[0052]
[0053] Among them, M att represents 1×1 convolution, represents a 1×1 convolution with SiLU activation layer, so the local attention matrix W att It is learned based on query features and contextualized key features, that is, the "self-attention" of local areas is enhanced by mining contextual features;
[0054] Step 5-3-4: Summarize the value feature matrix V and perform Softmax operation on the channel dimension to calculate the dynamic context self-attention weight matrix, which is expressed as follows:
[0055]
[0056] in Represents the attention score of the Softmax operation;
[0057] Step 5-3-5: Convert the static context feature K Static and dynamic context feature K dynamic The fusion is performed through the channel superposition fusion mechanism; in this way, P i,40×40×128 and P i,20×20×256 Perform feature enhancement.
[0058] Preferably, step 6 includes the following sub-steps:
[0059] Step 6-1: Since the steel bars of the tunnel lining are usually arranged side by side, i The length of the steel bar target area in the detection result image is estimated as the distance between the leftmost detection box and the rightmost detection box. Assuming that the number of steel bar targets predicted in step 5-4 is N, the horizontal coordinate set of each target center is i is the sequence value from left to right, and the target area length of the reinforcement is calculated as:
[0060]
[0061] Step 6-2: Calculate the average distance between predicted targets Assuming that the preset spacing is d, and the threshold between the average spacing and the preset spacing is θ, the condition for determining the lack of reinforcement is:
[0062]
[0063] Step 6-3: Determine the current segmentation graph F according to the defect determination condition i The test result will indicate whether there is missing tendon. If there is missing tendon, the word “Lackrebar” will be marked.
[0064] The advantages of the present invention compared with the prior art are:
[0065] Compared with the traditional manual subjective identification method, the lightweight intelligent detection method of railway tunnel lining defects proposed in the present invention is based on deep learning to train a lining defect detection model, predicts the defect category and location of the inspection image, and improves the intelligence and efficiency of the identification.
[0066] Compared with machine learning methods, the proposed method can extract nonlinear features containing depth information based on convolutional neural networks. This feature is multi-scale, rotationally invariant and environmentally robust, which is beneficial to improving the detection accuracy of void structures and steel bars in complex scenarios.
[0067] Existing lining disease detection models are mostly aimed at hollow defects. The proposed method includes two types of defects, hollow structures and steel bars, in the network model training, and adds post-processing of steel bar targets for reinforcement deficiency judgment, so that the method of the present invention can effectively detect two types of defects, hollow lining and missing steel bars.
[0068] The method proposed in the present invention constructs a detection model for lining voids and steel bar targets. The detection model proposes a large number of innovative works based on the "end-to-end" detection architecture. First, in the feature extraction and fusion part, from the perspective of gradient diversion and redundant feature extraction, the lightweight convolution module is improved and constructed to ensure efficient feature representation capabilities and improve the reasoning speed. Secondly, a cross-layer weighted channel feature fusion method is proposed in the feature fusion network. By establishing a weight competition mechanism for fused features, the channel feature fusion effect is improved and the model parameters are greatly reduced. Finally, a feature enhancement structure is added to the model. This structure constructs a multi-head attention convolution module based on the context Transformer, which is used to enhance the feature interaction of the context and improve the detection performance in complex scenarios.
[0069] From the perspective of feasibility, the method of the present invention implements the process of "segmentation-detection-merging" for long tunnel lining images. Compared with directly detecting long lining radar images, the detection efficiency can be greatly improved. The lightweight detection model constructed by the proposed method has lower requirements for hardware memory and video memory, and thus has strong portability. In short, this method provides a lightweight and intelligent disease detection method for railway tunnel lining detection that replaces manual identification and has engineering practicality, which is conducive to realizing efficient and intelligent railway tunnel engineering. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0071] Figure 1 The overall structural framework of railway tunnel lining defect detection provided by the embodiment of the present invention;
[0072] Figure 2 A schematic diagram of sliding window segmentation of raw data provided by an embodiment of the present invention;
[0073] Figure 3 A network model diagram for lightweight detection of lining defects constructed according to an embodiment of the present invention;
[0074] Figure 4 Schematic diagram of an efficient layer aggregation feature extraction module provided by an embodiment of the present invention;
[0075] Figure 5 A flow chart of a feature enhancement module provided by an embodiment of the present invention;
[0076] Figure 6A characteristic decoupling network structure diagram provided by an embodiment of the present invention;
[0077] Figure 7 This is the result of merging radar images of long tunnel lining. DETAILED DESCRIPTION
[0078] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.
[0079] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0080] The present invention is further described in detail below in conjunction with the accompanying drawings.
[0081] Combination Figure 1 to Figure 7 The lightweight intelligent detection method for railway tunnel lining defects of the present invention adopts WINDOWS10 as the operating system, Intel i7-10875H as the processor, the main frequency is 2.90GHz, the memory is 16.00GB, and the experiments involved are debugged on the PyCharm software platform using Python and Pytorch architecture.
[0082] Along Figure 1 The overall structural framework and other specific sample diagrams are used to illustrate the method in this paper. The method includes the following steps:
[0083] Step 1: Acquire multiple radar images of railway tunnel linings of a specific length, which contain reflection structures of multiple types of structures such as secondary lining, steel bars, and primary supports;
[0084] Step 2: Based on the mileage information of each input data, the lining radar image is labeled, segmented, and the relative coordinates of the segmented image target are transformed to construct an image dataset for training. The specific operation method is as follows: Figure 2 As shown;
[0085] Step 2-1: The segmentation method adopts sliding window segmentation, and the step size s of the sliding window is set to 1000 pixels, the corresponding segmentation image mileage is about 10m, and the segmentation margin g is 200 pixels;
[0086] Step 2-2: To prevent the segmentation operation from damaging the steel bars and the empty targets and to ensure the integrity of the targets, the segmentation edge is segmented and shifted. It is determined based on whether there is a target at each window edge. If the segmentation line passes through the target, the shifting process is implemented. If the segmentation line does not pass through the target, the segmentation area width remains unchanged. In this way, several segmentation maps are obtained;
[0087] Step 2-3: For the convenience of training, the initial marked target box position is converted into the target box position coordinates in the segmentation map. Assume that the i-th segmentation map F i The original parameters of the target box in are (x1, y1, x2, y2) and the delayed pixels are e i , then the relative position (x1', y1', x2', y2') is calculated as follows:
[0088]
[0089] y′1=y1
[0090]
[0091] y′2=y2
[0092] Step 3: Perform preprocessing operations such as adaptive scaling and padding on the segmented image set of the original data to obtain a preprocessed image with a fixed scale of 640×640 pixels; that is, for each segmentation image F i Scale normalization is performed, and adaptive scaling and padding are used to obtain a fixed input scale of c·c (c=640). Assume that the segmentation map F i The size of is a×b (a≥b), and the adaptive scaling factor is defined as the inverse of the ratio of the long side a to the fixed length c, that is, After scaling, the long side is scaled to c, and the short side length is λ·b. To convert to a normalized scale, the short sides on both sides are filled with 0, and the calculation formula for the single-side short side filling pixel size m is:
[0093]
[0094] Step 4: Construct a tunnel lining structure detection network, train a detection model for steel bars and hollow structures based on the training set target box annotation data, and obtain model parameters and neural network weight files; that is, construct a convolutional neural network for hollow structure and steel bar recognition, such as Figure 3 As shown, it includes an efficient feature extraction network, a lightweight feature fusion network, a feature enhancement module and a feature decoding structure. The positions and category labels of all segmentation map targets are input to train the neural network to obtain weight parameters;
[0095] Step 5: Perform preprocessing operations such as steps 2 and 3 on the image to be inspected. Based on the training weight parameters of step 4, use Figure 3 The network structure shown is used for reasoning to detect the categories and location parameters of the steel bars and void structures in all segmentation images;
[0096] Step 5-1: Input the preprocessed image into the efficient feature extraction network to obtain the initial feature maps F at three scales i,80×80×128 、F i,40×40×256 、F i,20×20×512 ;
[0097] Step 5-1-1: Use deep separable convolution, batch normalization, and SiLU function activation processing (DW-CBS module) to perform shallow feature extraction on the input image;
[0098] Step 5-1-2: Then, GhostELA-DWC3 module and DW-CBS module are stacked multiple times to perform multi-scale feature extraction. GhostELA-DWC3 is an improved deep feature extraction module, and its structure is as follows: Figure 4 Specifically, the DW-CBS module and the GhostELA-DWC3 module are superimposed twice to obtain the feature map F i,80×80×128 , superimposed three times to obtain the feature map F i,40×40×256 , stack four times to get F i,20×20×512 ;
[0099] Step 5-2: Use a lightweight feature fusion network to perform cross-layer weighted fusion of the initial features to obtain a fusion feature map P of three scales i,80×80×64 , P i,40×40×128 , P i,20×20×256 ;
[0100] Step 5-2-1: F i,20×20×512 The SPPF block is used for spatial pyramid pooling, and then GhostELA-DWC3, DW-CBS modules and nearest neighbor interpolation are used to obtain the double upsampled feature F. i,40×40×128 , this feature is consistent with F i,40×40×256 Splice in the channel direction to obtain the upsampled fusion feature map F i,40×40×384 In this way, another scale of upsampled fusion feature map F is obtained i,80×80×192 ;
[0101] Step 5-2-2: Feature map F i,80×80×192 The GhostELA-DWC3 module is further used to obtain the fusion feature map P i,80×80×64 Based on the idea of cross-layer weighted fusion, the channel weight W of the multi-layer feature map is calculated by fast normalization:
[0102]
[0103] where w i is the learnable weight, and ε = 0.0001 is used for stable propagation of gradients. i Then apply SiLU activation function to ensure the weighted value w i > 0. Then, each normalized weight value is limited to between 0 and 1.
[0104] Step 5-2-3: Initial feature map, upsampled fusion feature map F i,80×80×128 And the fusion feature map P i,80×80×64 Perform weighted splicing in the channel direction to obtain the fusion feature map P i,40×40×128 In this way, the fusion feature map P is obtained i,20×20×256 ;
[0105] Step 5-3: Further, adopt Figure 5 The fusion feature map P i,80×80×64 , P i,40×40×128 , P i,20×20×256 Implement context feature enhancement to obtain enhanced feature map
[0106] Step 5-3-1: Fusion feature map P i,80×80×64 Take it as an example, and convert it into a query Q, key K, and value V vector. The formula is as follows:
[0107] Q=P i,80×80×64 ·M q
[0108]
[0109] V=P i,80×80×64 ·M v
[0110] Among them, M q , M v are the embedding matrices used to transform query, key, and value vectors, respectively.
[0111] Step 5-3-2: Assume the center key of the context area is X cen , the size of the surrounding area is ×(=3), then × times of convolution can be calculated to obtain the key vector information of each surrounding area. Similar to the sliding window convolution, the learned context key K Static Reflects the static information of the center and surroundings.
[0112] Step 5-3-3: Concatenate the context key and query vector to form a synthetic key [K Static ,]. Self-attention encoding is performed by using two consecutive 1×1 convolutions:
[0113]
[0114] Among them, M att represents 1×1 convolution, represents a 1×1 convolution with SiLU activation layer. Therefore, the local attention matrix W att It is learned based on query features and contextualized key features, that is, the "self-attention" of local areas is enhanced by mining contextual features;
[0115] Step 5-3-4: Summarize the value feature matrix (V) and perform Softmax operation on the channel dimension to calculate the dynamic context self-attention weight matrix, which is expressed as follows:
[0116]
[0117] in Represents the attention score of the Softmax operation;
[0118] Step 5-3-5: Convert the static context feature K Static and dynamic context feature K dynamic The fusion is performed through the channel superposition fusion mechanism. In this way, P i,40×40×128 and P i,20×20×256 Perform feature enhancement.
[0119] Step 5-4: Enhance the feature map Input feature decoding structure to obtain prediction vector. The specific process is as follows Figure 6 As shown. The final prediction vector contains (t x ,t y ,t w ,t h ), Sorce representing the confidence score of the current prediction box confi And Probability representing the probability of the target category class To normalize the output of the predicted target box, the coordinates of the center point of the predicted box (b x ,b y ,b w ,b h ) is calculated from the relative predicted coordinates of the corresponding grid (t x ,t y ,t w ,t h ) indicates that the calculation formula is as follows:
[0120] b x =c x +2σ(t x )-0.5
[0121] b y =c y+2σ(t y )-0.5
[0122] b w =p w ·(2σ(t w )) 2
[0123] b h =p h ·(2σ(t h )) 2
[0124] Where (c x ,c y ) is the relative coordinate of the upper left corner of the grid responsible for predicting the target.
[0125] Step 6: Determine segmentation diagram F based on preset spacing of steel bars and number of detected steel bars i The test results are used to determine whether there is a lack of reinforcement and obtain the post-processing test result diagram;
[0126] Step 6-1: Since the steel bars of the tunnel lining are usually arranged side by side, i The length of the steel bar target area in the detection result image can be estimated as the distance between the leftmost detection box and the rightmost detection box. Assuming that the number of steel bar targets predicted in step 5-4 is N, the horizontal coordinate set of each target center is i is the sequence value from left to right, and the target area length of the reinforcement is calculated as:
[0127]
[0128] Step 6-2: Calculate the average distance between predicted targets Assuming that the preset spacing is d, and the threshold between the average spacing and the preset spacing is θ, the defect determination condition is: The formula is:
[0129]
[0130] Step 6-3: Determine the current segmentation graph F according to the defect determination condition i The test result shows whether there is a lack of rebar. If there is a lack of rebar, the word "Lack rebar" will be marked;
[0131] Step 7: Merge all post-processing detection segmentation maps, that is, splice all segmentation map detection results based on the segmentation parameters of step 2 to obtain the final detection result of the long detection with the same size as the input data. The detection result is as follows: Figure 7 shown.
[0132] The present invention and its embodiments are described above, and such description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if ordinary technicians in the field are inspired by it, without departing from the purpose of the invention, they can design a structure and embodiment similar to the technical solution without creativity, which should belong to the protection scope of the present invention.
Claims
1. A lightweight intelligent detection method for railway tunnel lining defects, characterized in that: The following steps are involved: Step 1: Acquire multiple radar images of railway tunnel linings of a specific length, which contain multiple types of reflective structures such as secondary lining, steel bars, and primary supports; Step 2: According to the mileage information of each image input data, the lining radar image is labeled, segmented, and the relative coordinates of the segmented image targets are transformed to construct an image dataset for training; Step 3: For each segmentation map F i Scale normalization is performed, and adaptive scaling and padding are used to obtain a fixed input scale of c·c, where c = 640; assuming that the segmentation map F i The size of is a×b (a≥b), and the adaptive scaling factor is defined as the inverse of the ratio of the long side a to the fixed length c, that is, After scaling, the long side is scaled to c, and the short side length is λ·b. To convert to a normalized scale, the short sides on both sides are filled with 0, and the calculation formula for the single-side short side filling pixel size m is: Step 4: Construct a convolutional neural network for hollow structure and steel bar identification, including an efficient feature extraction network, a lightweight feature fusion network, a feature enhancement module, and a feature decoding structure; input the location and category labels of all segmentation map targets to train the neural network to obtain weight parameters; Step 5: Perform the preprocessing operations of steps 2 and 3 on the image to be inspected, and then based on the training weight parameters of step 4 and the network structure inference, detect the categories and location parameters of the steel bars and the hollow structures in all the segmented images; Step 6: According to the segmentation map F i The test results can be used to determine whether there is a lack of tendons; Step 7: Based on the segmentation parameters of step 2, all segmentation map detection results are spliced to obtain a long detection result of the same size as the input data.
2. A lightweight intelligent detection method for railway tunnel lining defects according to claim 1, characterized in that: The step 2 includes the following sub-steps: Step 2-1: The segmentation method adopts sliding window segmentation, and the step size s of the sliding window is set to 1000 pixels, the corresponding segmentation image mileage is about 10m, and the segmentation margin g is 200 pixels; Step 2-2: Perform segmentation and shifting processing on the segmentation edge to ensure the integrity of the target and prevent the segmentation operation from damaging the steel bars and the empty targets; determine whether there is a target at each window edge, if the segmentation line passes through the target, perform shifting processing, if the segmentation line does not pass through the target, the segmentation area width remains unchanged; in this way, obtain several segmentation maps; Step 2-3: To facilitate training, convert the initial marked target box position into the target box position coordinates in the segmentation map; assuming that the i-th segmentation map F i The original parameters of the target box in are (x1, y1, x2, y2) and the delayed pixels are e i , then the relative position (x1', y1', x2', y2') is calculated as follows: y′1=y1 y′2=y2。 3. The lightweight intelligent detection method for railway tunnel lining defects according to claim 1 is characterized by: The step 5 comprises the following sub-steps: Step 5-1: Input the preprocessed image into the efficient feature extraction network to obtain the initial feature maps F at three scales i,80×80×128 、F i,40×40×256 、F i,20×20×512 ; Step 5-2: Use a lightweight feature fusion network to perform cross-layer weighted fusion of the initial features to obtain a fusion feature map P of three scales i,80×80×64 , P i,40×40×128 , P i,20×20×256 ; Step 5-3: Fusion feature map P i,80×80×64 , P i,40×40×128 , P i,20×20×256 Implement context feature enhancement to obtain enhanced feature map Step 5-4: Enhance the feature map Input feature decoding structure to obtain prediction vector; finally, the prediction vector contains (t x ,t y ,t w ,t h ), Sorce representing the confidence score of the current prediction box confi And Probability representing the probability of the target category class ; To normalize the output of the predicted target box, the coordinates of the center point of the predicted box (b x ,b y ,b w ,b h ) is calculated from the relative predicted coordinates of the corresponding grid (t x ,t y ,t w ,t h ) indicates that the calculation formula is as follows: b x =c x +2σ(t x )-0.5 b y =c y +2σ(t y )-0.5 b w =p w ·(2σ(t w )) 2 b h =p h ·(2σ(t h )) 2 ; Where (c x ,c y ) is the relative coordinate of the upper left corner of the grid responsible for predicting the target.
4. A lightweight intelligent detection method for railway tunnel lining defects according to claim 3, characterized in that: The step 5-1 includes the following sub-steps: Step 5-1-1: Use depth-separable convolution, batch normalization, and SiLU function activation processing on the input image, and use the DW-CBS module to perform shallow feature extraction; Step 5-1-2: Multiple stacking of GhostELA-DWC3 module and DW-CBS module for multi-scale feature extraction, GhostELA-DWC3 is an improved deep feature extraction module; the stacking method is: stacking DW-CBS module and GhostELA-DWC3 module twice to obtain feature map F i,80×80×128 , superimposed three times to obtain the feature map F i,40×40×256 , stack four times to get F i,20×20×512 .
5. The lightweight intelligent detection method for railway tunnel lining defects according to claim 3 is characterized by: The step 5-2 includes the following sub-steps: Step 5-2-1: F i,20×20×512 The SPPF block is used for spatial pyramid pooling, and then GhostELA-DWC3, DW-CBS modules and nearest neighbor interpolation are used to obtain the double upsampled feature F. i,40×40×128 , this feature is consistent with F i,40×40×256 Splice in the channel direction to obtain the upsampled fusion feature map F i,40×40×384 ; In this way, another scale of upsampled fusion feature map F is obtained i,80×80×192 ; Step 5-2-2: Feature map F i,80×80×192 The GhostELA-DWC3 module is further used to obtain the fusion feature map P i,80×80×64 ; Based on the idea of cross-layer weighted fusion, the channel weight W of the multi-layer feature map is calculated by fast normalization: Where w i is the learnable weight, ε = 0.0001 is used for stable propagation of gradients; at each w i Then apply SiLU activation function to ensure the weighted value w i >0; then, each normalized weight value is limited to between 0 and 1; Step 5-2-3: Initial feature map, upsampled fusion feature map F i,80×80×128 And the fusion feature map P i,80×80×64 Perform weighted splicing in the channel direction to obtain the fusion feature map P i,40×40×128 ; In this way, the fusion feature map P is obtained i,20×20×256 .
6. A lightweight intelligent detection method for railway tunnel lining defects according to claim 3, characterized in that: The step 5-3 includes the following sub-steps: Step 5-3-1: Fusion feature map P i,80×80×64 Transformed into query Q, key K and value V vectors, the formula is as follows: Q=P i,80×80×64 ·M q V=P i,80×80×64 ·M v ; Among them, M q , M v are the embedding matrices used to transform query, key, and value vectors, respectively; Step 5-3-2: Assume the center key of the context area is X cen , the size of the surrounding area is k×k, where k=3, then k×k convolutions are calculated to obtain the key vector information of each surrounding area; the learned context key K Static Reflects the static information of the center and surroundings; Step 5-3-3: Concatenate the context key and query vector to form a synthetic key [K Static ,Q]; self-attention encoding is performed by using two consecutive 1×1 convolutions: Among them, M att represents 1×1 convolution, represents a 1×1 convolution with SiLU activation layer, so the local attention matrix W att It is learned based on query features and contextualized key features, that is, the "self-attention" of local areas is enhanced by mining contextual features; Step 5-3-4: Summarize the value feature matrix V and perform Softmax operation on the channel dimension to calculate the dynamic context self-attention weight matrix, which is expressed as follows: in Represents the attention score of the Softmax operation; Step 5-3-5: Convert the static context feature K Static and dynamic context feature K dynamic The fusion is performed through the channel superposition fusion mechanism; in this way, P i,40×40×128 and P i,20×20×256 Perform feature enhancement.
7. The lightweight intelligent detection method for railway tunnel lining defects according to claim 1 is characterized by: The step 6 comprises the following sub-steps: Step 6-1: Since the steel bars of the tunnel lining are usually arranged side by side, i The length of the steel bar target area in the detection result image is estimated as the distance between the leftmost detection box and the rightmost detection box. Assuming that the number of steel bar targets predicted in step 5-4 is N, the horizontal coordinate set of each target center is i is the sequence value from left to right, and the target area length of the reinforcement is calculated as: Step 6-2: Calculate the average distance between predicted targets Assuming that the preset spacing is d, and the threshold between the average spacing and the preset spacing is θ, the condition for determining the lack of reinforcement is: Step 6-3: Determine the current segmentation graph F according to the defect determination condition i The test result will indicate whether there is missing ribs. If there is missing ribs, the word "Lackrebar" will be marked.
Citation Information
Patent Citations
In-service tunnel lining structure defect identification method and system based on deep learning
CN114548278A
Tunnel lining geological radar data intelligent identification method
CN116340770A