Real-time Semantic Segmentation Tunnel Overbreak and Underbreak Monitoring Method and System

Through the FP-Former semantic segmentation deep neural network, the real-time and accuracy problems of ultra-under-digging monitoring in tunnel construction are solved, and efficient and accurate tunnel ultra-under-digging detection is achieved to ensure construction safety and quality.

CN119649022BActive Publication Date: 2025-07-04EAST CHINA JIAOTONG UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411621008.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-07-04
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

The lack of real-time and accurate over-excavation monitoring methods in existing tunnel construction, resulting in increased construction safety risks and costs. Traditional methods rely on manual operations and lack of real-time segmentation accuracy and speed.

Method used

The FP-Former semantic segmentation deep neural network is used to monitor tunnel hyper-under-digging area through data cleaning and enhancement, and combine cross-attention and deep aggregation pyramid pool module to achieve high-precision and fast tunnel hyper-under-digging detection.

Benefits of technology

High-precision and rapid detection of tunnel over-under-excavation areas are achieved, construction safety and quality are ensured, real-time feedback mechanism is provided, safety risks are reduced and construction efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119649022B_ABST
    Figure CN119649022B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for real-time semantic segmentation of tunnel overbreak and underbreak monitoring, including: obtaining high-resolution images of different excavation surface areas during tunnel construction suitable for excavation under different excavation conditions and geological backgrounds, and labeling overbreak and underbreak area tags; data cleaning and enhancement, obtaining a number of tunnel overbreak and underbreak images with a resolution of 512×1024 after data cleaning and enhancement processing to form a target data set; constructing an FP-Former semantic segmentation deep neural network, training the FP-Former semantic segmentation deep neural network using the images in the target data set to monitor overbreak and underbreak areas and output a prediction probability map; performing binarization processing on the output prediction probability map, marking pixels with a probability greater than 0.5 as overbreak areas, and vice versa as underbreak areas. The present invention can effectively extract features of different scales, thereby improving the segmentation accuracy. When running on GPU-class devices, it shows a good balance of performance and efficiency, making large-scale data processing possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent construction technology, and in particular to a real-time semantic segmentation tunnel over-excavation and under-excavation monitoring method and system. Background Art

[0002] Tunnel engineering plays a vital role in modern infrastructure construction, especially in urban transportation, subway systems and water conservancy projects. However, in the actual construction process, over-excavation and under-excavation often occur. Traditional tunnel excavation methods rely on manual operation and are limited by the experience and judgment of construction workers. When faced with emergencies, construction workers may make wrong decisions, resulting in over-excavation or under-excavation. This phenomenon may not only lead to safety hazards of tunnel structures, but also increase the complexity and cost of subsequent maintenance.

[0003] Overbreak usually refers to the situation where the actual amount of earth excavated exceeds the design requirements during the excavation process. This situation may lead to a decrease in the stability of the tunnel wall and increase the risk of collapse. Underbreak means failure to excavate to the predetermined depth or width, which may lead to insufficient tunnel capacity and affect subsequent equipment installation and operational safety. During the construction process, if there is a lack of effective real-time monitoring methods, the construction team may not be able to understand the changes in the excavation progress and status in a timely manner, and miss the best time to correct the problem, which increases the risk of overbreak and underbreak. Therefore, in order to ensure the overall quality and safety of tunnel projects, real-time and accurate monitoring of the state of the excavation surface and identification of overbreak and underbreak have become the key to ensuring the quality of tunnel construction.

[0004] In recent years, deep learning and computer vision technologies have made remarkable progress in various fields, especially in image processing and analysis. Among them, semantic segmentation, as an important image processing task, has been widely studied and applied.

[0005] Real-time semantic segmentation technology provides strong support for tunnel construction monitoring, and can quickly and accurately analyze images obtained in the tunnel. Through deep learning algorithms, the system can automatically identify over-excavation and under-excavation areas in the image and clearly mark potential construction problems. This precise recognition capability allows construction personnel to obtain feedback in the first time and understand the construction progress and safety status in a timely manner. Once the system detects over-excavation or under-excavation, construction personnel can quickly take corresponding corrective measures, thereby significantly reducing safety risks. This not only ensures the safety of construction personnel, but also helps maintain the structural integrity of the tunnel and avoid greater losses caused by the expansion of problems. In addition, the real-time monitoring and feedback mechanism improves construction efficiency, allowing construction teams to respond quickly in a dynamically changing environment and optimize construction plans and resource allocation.

[0006] Existing semantic segmentation uses convolutional neural networks (CNNs) and their various variants, such as U-Net and DeepLab. Although good results have been achieved, due to their large number of parameters and slow processing speed, they cannot be well deployed on mobile devices, and the accuracy and speed of real-time segmentation need to be improved. Based on the above problems, the present invention proposes a real-time semantic segmentation tunnel overbreak and underbreak monitoring method that can be quantified more quickly and accurately, which not only improves the accuracy of real-time detection of tunnel overbreak and underbreak but also provides the fastest feedback, enabling engineers to deeply understand the health status of the tunnel, timely identify potential problems, and thus implement more effective maintenance strategies in a timely manner. Summary of the Invention

[0007] The object of the present invention is to overcome the above-mentioned technical drawbacks and provide an intelligent real-time semantic segmentation tunnel overbreak and underbreak monitoring method and system that can monitor the area of overexcavated and under-excavated soil in the tunnel in real time.

[0008] To achieve the above object, the technical solution of the present invention is as follows:

[0009] In the first aspect, the present invention provides a real-time semantic segmentation tunnel overbreak and underbreak monitoring method, characterized in that the monitoring method includes the following contents:

[0010] Step 1: Obtain high-resolution images of different soil excavation areas during tunnel construction suitable for excavation under different excavation conditions and geological backgrounds, and label the overexcavation and under-excavation area labels.

[0011] Step 2: Data cleaning and enhancement

[0012] 2.1 Data cleaning:

[0013] ① Mark the images with a sharpness variance lower than 10 as blurred images, and perform deletion operations, retaining the images with a sharpness variance not lower than 10.

[0014] ② Calculate the hash value of the image, delete the images with the same hash value, and only retain one image with the same hash value to ensure that each image has a unique hash value.

[0015] ③ Delete the invalid images containing useless information to ensure that the remaining images are all related to the excavation conditions and the overbreak and underbreak status. The invalid images include completely black screens, completely white screens, or other irrelevant backgrounds.

[0016] 2.2 Data enhancement is carried out in the following steps in sequence:

[0017] ① Rotation: Randomly rotate the image, and the rotation angle is between 30° and +30° to generate images with multiple angles to enhance the adaptability of the model to different directions.

[0018] ② Scaling: Randomly scale the image within the range of 50% - 200% to simulate the situation of shooting at different distances;

[0019] ③ Cropping: Randomly crop a part of the image and crop the image resolution to 512×1024;

[0020] ④ Color adjustment: Adjust the brightness, contrast, and saturation of the image to simulate images under different lighting conditions;

[0021] ⑥ Adding noise: Add random noise or simulate sensor noise to the image to improve the robustness of the model under low-quality images or interference conditions;

[0022] ⑦ Flipping: Randomly flip the image horizontally or vertically;

[0023] After data cleaning and enhancement processing, several tunnel overbreak and underbreak images with a resolution of 512×1024 are obtained to form the target dataset;

[0024] Step 3: Construct the FP-Former semantic segmentation deep neural network

[0025] The FP-Former semantic segmentation deep neural network includes a backbone block, a residual block, a deep double-resolution module, two stacked FPformer modules, a deep aggregation pyramid pooling module, and a pixel-level classification head; the backbone block is composed of two cascaded 3×3 convolutional layers, and each 3×3 convolutional layer includes 3×3 convolution, batch normalization operation, and RELU operation; the residual block is composed of four consecutive basic residual blocks in series, the input of the residual block is the output of the backbone block, and the output of the residual block is the high-resolution feature map N, and the high-resolution feature map N is dimension-reduced by stride convolution to obtain the low-resolution feature map n; the deep double-resolution module includes a high-resolution branch and a low-resolution branch, and the high-resolution feature map N and the low-resolution feature map n enter the high-resolution branch and the low-resolution branch of the deep double-resolution module respectively; the deep double-resolution module outputs the low-resolution feature f and the high-resolution feature F; then it enters the two stacked FPformer modules, and outputs the high-resolution feature O and the low-resolution feature o;

[0026] The low-resolution feature o enters the deep aggregation pyramid pooling module to extract context information, and the output feature information after processing by the deep aggregation pyramid pooling module is fused with the high-resolution feature O to obtain an output feature map with a stride of 8;

[0027] Finally, the output feature map with a stride of 8 is passed to the pixel-level classification head for predicting dense semantic labels and outputting a prediction probability map;

[0028] Step 4: Train the FP-Former semantic segmentation deep neural network using the images in the target dataset, monitor the over-excavation and under-excavation areas using the trained FP-Former semantic segmentation deep neural network, and output a prediction probability map; perform binarization processing on the output prediction probability map, set the threshold to 0.5, mark the pixels with a probability greater than 0.5 as over-excavation areas, and vice versa as under-excavation areas, and then obtain the areas of the over-excavation or under-excavation areas respectively for engineering monitoring.

[0029] Further, the pixel-level classification head consists of a 3×3 convolutional layer and a 1×1 convolutional layer, and both convolutional layers include batch normalization and ReLU activation functions.

[0030] Further, the deep double-resolution module includes a high-resolution branch, a low-resolution branch, and a bilateral fusion unit. The high-resolution branch includes two 3×3 convolutional operations and two batch normalization operations, and uses ReLU as the activation function; the high-resolution feature map N sequentially passes through a 3×3 convolutional operation, a batch normalization operation, a ReLU function, a 3×3 convolutional operation, and a batch normalization operation to extract feature information. At the same time, the high-resolution feature map N is subjected to a residual connection with the output result of the last batch normalization operation, and the feature after the residual connection is denoted as N1; the number of channels of both 3×3 convolutional operations is 128;

[0031] The low-resolution branch has the same structure as the high-resolution branch, also including two 3×3 convolutional operations, two batch normalizations, using ReLU as the activation function and residual connection. Only the number of channels of the two 3×3 convolutional operations is 256, and residual connection is also used. The input of the low-resolution branch is the low-resolution feature map n, and the feature after the residual connection is denoted as n1;

[0032] The bilateral fusion unit includes high-to-low fusion and low-to-high fusion. For high-to-low fusion, N1 is downsampled through a 3×3 convolutional sequence with a stride of 2 and 256 channels, and then after batch normalization processing, it is concatenated and fused with n1 to obtain n2;

[0033] For low-to-high fusion, n1 is first compressed through a 1×1 convolutional layer with 128 channels, and after batch normalization, it is upsampled through bilinear interpolation and then concatenated and fused with N1 to obtain N2;

[0034] n2 is re-fused with N2 processed by the ReLU activation function after passing through the ReLU activation function and upsampling processing to obtain a new high-resolution feature map F;

[0035] At the same time, n2 is processed by the ReLU activation function and subjected to stride convolution for dimensionality reduction to obtain a new low-resolution feature map f, and the stride of the stride convolution is 16 at this time.

[0036] Furthermore, the FPformer module has a low-resolution branch and a high-resolution branch. In the low-resolution branch, the new low-resolution feature map f passes through a GPU-friendly attention module to capture high-level global context. After the attention map output by the GPU-friendly attention module is feature-fused with the new low-resolution feature map f, the fused feature map enters a KAN module and then undergoes a 3×3 convolution operation to extract features, and then a residual connection is made with the input of the KAN module to obtain the feature map X l ;

[0037] In the high-resolution branch, a cross-attention module is used. The new high-resolution feature map F is concatenated and fused with the feature map X l to obtain the feature map X h , and the feature map X h is used as the input of the Q branch to the cross-attention module. At the same time, the feature map X l undergoes convolution and pooling operations to obtain the feature map X c , and the feature map X c is used as the input of the K and V branches to the cross-attention module. After being processed by the cross-attention module, the output attention map is then residually connected to the feature map X h , and then features are extracted through two stacked 3×3 convolutional layers, and a residual connection is made with the input of the two stacked 3×3 convolutional layers to obtain the feature map F1;

[0038] The feature map F1 is concatenated with the feature map X l to obtain the low-resolution feature o, which is used as the output of the low-resolution branch; the low-resolution feature o is then upsampled and feature-fused with the feature map F1 to generate the high-resolution feature O, which is used as the output of the high-resolution branch.

[0039] Furthermore, the GPU-friendly attention module is expressed as:

[0040]

[0041] where Δ = 0.05 ∈ (-0.1, 0.1) is a correction coefficient for adjusting the GDN output; α = 0.5 is a weight parameter for adjusting the influence of V g ; GFA represents the GPU-friendly attention operation; GDN is the grouped double normalization operation; X represents the input feature; is the learnable parameter in the GPU-friendly attention module; M g = M×H, M is the parameter dimension, d is the feature dimension, H is the number of heads of the multi-head attention mechanism, and T represents matrix transpose;

[0042] The cross-attention module is expressed as:

[0043]

[0044] Among them, β = 1 is a weight parameter used to adjust the influence of V c ; CA is the cross-attention module operation; X h , X l respectively represent the feature maps on the high-resolution branch and the low-resolution branch. is a set of matrix operations including splitting, permuting, and reshaping, and the inputs to the K and V branches obtained after processing the feature map X c are respectively denoted as K c , V c ; d h represents the feature dimension of the high-resolution branch. At the same time, the feature map X c is obtained by calculating X l through a function θ composed of a pooling layer and a convolutional layer; the spatial size of X c represents the number of tokens generated from X l .

[0045] Furthermore, the depth aggregation pyramid pooling module includes a pooling operation with a stride of 2 and a kernel of 5, a pooling operation with a stride of 4 and a kernel of 9, a pooling operation with a stride of 8 and a kernel of 17, and a global average pooling operation;

[0046] The outputs after the three pooling operations and the global average pooling operation are each processed through a 1×1 convolution and an upsampling operation, resulting in a total of four branches; at the same time, the low-resolution feature o is processed through a 1×1 convolution to obtain an output y1, and y1 and the output of the first branch are processed through a dilated convolution with a dilation rate of 4 to obtain an output y2, y2 and the output of the second branch are processed through a dilated convolution with a dilation rate of 8 to obtain an output y3, y3 and the output of the third branch are processed through a dilated convolution with a dilation rate of 12 to obtain an output y4, and y4 and the output of the fourth branch are processed through a dilated convolution with a dilation rate of 16 to obtain an output y5; after y1, y2, y3, y4, and y5 are concatenated, they are processed through a 1×1 convolution and then concatenated with y1 to obtain the output of the depth aggregation pyramid pooling module.

[0047] Furthermore, the FP-Former semantic segmentation deep neural network is trained using the Adam W optimizer. During training, the initial learning rate is 0.0004, and the weight decay is 0.0125; a poly learning strategy with a power of 0.9 is adopted to reduce the learning rate, and 484 epochs and a batch size of 12 are set.

[0048] Further, calculate the ratio of the over-excavated or under-excavated area to the total tunnel area to obtain the area percentage, and set an area percentage threshold. When the detected area percentage of over-excavation or under-excavation exceeds the set area percentage threshold, an alarm is immediately triggered. Different levels of alarms are set according to the severity of the over-excavated or under-excavated area.

[0049] ① Warning level: Slight over-excavation or under-excavation, the excavation posture of the excavator needs to be corrected.

[0050] ② Alarm level: Moderate over-excavation or moderate under-excavation, immediate measures need to be taken to re-excavate the under-excavated part and adjust the excavation posture.

[0051] ③ Emergency level: Severe over-excavation or severe under-excavation, the operation needs to be stopped and an in-depth investigation is carried out, the excavation plan is re-set, and the deviation is corrected.

[0052] If it does not exceed the area percentage threshold range, the tunnel quality is qualified; otherwise, the quality is unqualified and an alarm is given and feedback is provided.

[0053] Further, check the area percentage. If the over-excavation area percentage is within [3%, 5%] of the total area or the under-excavation area percentage is within [2% - 3%] of the total area, it is set as the warning level; if the over-excavation area percentage is within (5% - 10%] of the total area or the under-excavation area percentage is within (3% - 5%] of the total area, it is set as the alarm level; if the over-excavation area percentage exceeds 10% of the total area or the under-excavation area percentage exceeds 5% of the total area, it is set as the emergency level; if the remaining area ratios are not within the above ranges, they are regarded as qualified.

[0054] In a second aspect, the present invention provides a real-time semantic segmentation tunnel over-excavation and under-excavation monitoring system, which includes the following modules:

[0055] Data collection module: Use drones, laser scanners or high-definition camera devices to obtain high-resolution images during tunnel construction, ensure coverage of different excavation surface areas, and record the real-time excavation depth and soil properties in combination with soil sensors and pressure sensors.

[0056] Data annotation module: Annotate the collected images to clearly distinguish the over-excavation and under-excavation image areas and make accurate labels.

[0057] Data cleaning and enhancement module: Remove blurred, duplicate or invalid images, and at the same time perform image preprocessing using operations such as rotation, scaling, cropping, color adjustment, adding noise and flipping.

[0058] Image Segmentation Module: Establish an FP-Former semantic segmentation deep neural network for image segmentation, connect it with the data annotation module and the data cleaning and enhancement module to obtain the processed image, and use this to train the FP-Former semantic segmentation deep neural network for detecting the prediction probability map of the over-excavation and under-excavation areas of the tunnel;

[0059] Over-Excavation and Under-Excavation Area Measurement Module: Binarize the prediction probability map, set the threshold to 0.5, mark the pixels with a probability greater than 0.5 as the over-excavation area, and vice versa as the under-excavation area. Calculate the area of the over-excavation and under-excavation areas, take the ratio of the over-excavation or under-excavation area to the total area of the tunnel to obtain the area percentage, and set the area percentage threshold;

[0060] Early Warning and Feedback Module: Connect with the over-excavation and under-excavation area measurement module, and according to the detected area percentage, promptly give an early warning to the background system and feedback the degree of over-excavation and under-excavation of the tunnel. The background system automatically issues an alarm and prompts relevant personnel to take measures.

[0061] Compared with the prior art, the beneficial effects of the present invention are:

[0062] In view of the deficiencies in the monitoring accuracy and real-time performance of tunnel over-excavation and under-excavation, the present invention proposes to use an FP-Former semantic segmentation deep neural network for image semantic segmentation. The FPformer module combines the advantages of the feature pyramid network and can effectively extract features at different scales, thereby improving the segmentation accuracy. When the method of the present invention runs on GPU-class devices, it shows a good balance between performance and efficiency, making large-scale data processing possible.

[0063] The method of the present invention fully considers the utilization of global context information. By using the cross-attention mechanism, the network can combine global information while analyzing local features, enhancing the recognition ability of tunnel over-excavation and under-excavation areas, allowing the network to dynamically focus on important features when processing images, thereby improving the accuracy of semantic segmentation.

[0064] In practical applications, especially in the tunnel construction environment, it can quickly and accurately detect the over-excavation and under-excavation areas, ensuring efficient real-time monitoring. Through high-precision quantification of these areas and analysis of the area percentage of the over-excavation area, construction personnel can obtain timely feedback and take corresponding measures to ensure the safety and quality of tunnel construction. The method of the present invention not only improves the monitoring efficiency but also provides an effective technical means for tunnel quality assurance and can play an important role in complex construction environments. Brief Description of the Drawings

[0065] Figure 1 It is a schematic structural diagram of the FP-Former semantic segmentation deep neural network in the present invention;

[0066] Figure 2 It is a schematic structural diagram of the depth double-resolution module in the FP-Former semantic segmentation deep neural network of the present invention;

[0067] Figure 3 It is a schematic structural diagram of the FPformer module in the FP-Former semantic segmentation deep neural network of the present invention;

[0068] Figure 4 It is a schematic structural diagram of the depth aggregation pyramid pooling module in the FP-Former semantic segmentation deep neural network of the present invention;

[0069] Figure 5 It is a schematic process diagram of the real-time semantic segmentation tunnel overbreak and underbreak monitoring system of the present invention. Specific implementation manners

[0070] Compared with the prior art, the present invention adopts the FP-Former semantic segmentation deep neural network, uses the FPformer module, and achieves a good compromise between performance and efficiency on GPU-class devices for semantic segmentation tasks; at the same time, without sacrificing efficiency, it makes full use of the global context and deeply utilizes attention to improve semantic segmentation, realizing high-precision and fast monitoring of the overbreak and underbreak area regions of the tunnel, being able to accurately and efficiently detect the overbreak and underbreak categories and quantify them, providing an effective means for ensuring the quality of the tunnel.

[0071] The steps of a real-time semantic segmentation tunnel overbreak and underbreak monitoring method of the present invention are as follows:

[0072] I. Dataset preparation

[0073] 1. Sensor data: Combine devices such as soil sensors and pressure sensors to record the real-time excavation depth and soil properties. The type, density, and permeability of the soil are closely related to tunnel operations. Record the soil properties in real time to ensure whether the selected area is suitable for excavation.

[0074] 2. Image acquisition: Use industrial cameras or high-definition imaging devices to obtain high-resolution images during tunnel construction, ensuring that different excavation surface areas can be collected.

[0075] II. Data annotation

[0076] 1. Overbreak and underbreak annotation: Manually annotate the overexcavation and underexcavation areas using image annotation tools (such as LabelMe, VGG Image Annotator, etc.), ensuring that each image has an accurate label.

[0077] 2. Diversity consideration: Ensure that the annotated samples cover different excavation conditions and geological backgrounds to improve the generalization ability of the model.

[0078] III. Data Cleaning and Augmentation

[0079] 1. Data Cleaning:

[0080] ① Use an image sharpness metric (such as calculating the Laplacian variance) to evaluate the sharpness of each image. Images with a sharpness variance below 10 will be marked as blurred images and then deleted.

[0081] ② Calculate the hash value of the image to detect duplicate images. The same hash value means these images are duplicates, and then the duplicate images are deleted to ensure that each image in the dataset is unique.

[0082] ③ Check if the image contains useless information. For example, images with a completely black screen, white screen, or other irrelevant backgrounds are invalid images. Then these invalid images are deleted to ensure that the remaining images are all related to the excavation conditions and over - under - excavation status.

[0083] Through the above operations, blurred, duplicate, or invalid images are removed to ensure the quality of the dataset.

[0084] 2 Data augmentation is carried out in the following steps in sequence:

[0085] ① Rotation: Randomly rotate the image (rotation angle between 30° and +30°) to generate images at multiple angles to enhance the model's adaptability to different directions.

[0086] ② Scaling: Randomly scale the image (scaling range between 50% - 200%) to simulate the situation of being taken at different distances.

[0087] ③ Cropping: Randomly crop a part of the image and crop the image resolution to 512×1024.

[0088] ④ Color adjustment: Adjust the brightness, contrast, saturation, etc. of the image to simulate images under different lighting conditions.

[0089] ⑥ Adding noise: Add random noise (such as Gaussian noise) or simulate sensor noise to the image to improve the model's robustness under low - quality images or interference conditions.

[0090] ⑦ Flipping: Randomly flip the image horizontally or vertically.

[0091] Through data cleaning and augmentation, the quality and diversity of the dataset can be effectively improved, providing a more reliable and rich sample basis for the subsequent training of the model, thereby enhancing the performance and effect of the model in practical applications.

[0092] IV. Dataset Division

[0093] After the above images are subjected to data cleaning and enhancement processing, 10,000 tunnel overbreak and underbreak pictures with a resolution of 512×1024 are obtained.

[0094] Construct the target dataset R for training the FP-Former semantic segmentation deep neural network with the images after data cleaning and enhancement processing. The target dataset R is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1 for the training and validation of the FP-Former semantic segmentation deep neural network.

[0095] V. Construct the FP-Former semantic segmentation deep neural network

[0096] The FP-Former semantic segmentation deep neural network, as Figure 1 shown, includes a backbone block, a residual block, a deep double-resolution module, two stacked FPformer modules, a deep aggregated pyramid pooling module, and a pixel-level classification head; the backbone block is composed of two 3×3 convolutional layers in series, and each 3×3 convolutional layer includes a 3×3 convolution, a batch normalization operation, and a RELU operation; the residual block is composed of four consecutive basic residual blocks in series, the input of the residual block is the output of the backbone block, and the output of the residual block is a high-resolution feature map N, and the high-resolution feature map N is dimension-reduced by a stride convolution to obtain a low-resolution feature map n; the deep double-resolution module includes a high-resolution branch and a low-resolution branch, and the high-resolution feature map N and the low-resolution feature map n enter the high-resolution branch and the low-resolution branch of the deep double-resolution module respectively; the deep double-resolution module outputs a low-resolution feature f and a high-resolution feature F; then it enters two stacked FPformer modules, and outputs a high-resolution feature O and a low-resolution feature o;

[0097] The low-resolution feature o enters the deep aggregated pyramid pooling module to extract context information, and the output feature information after the deep aggregated pyramid pooling module is processed is fused with the high-resolution feature O to obtain an output feature map with a stride of 8;

[0098] Finally, the output feature map with a stride of 8 is passed to the pixel-level classification head for predicting dense semantic labels and outputting a prediction probability map.

[0099] In this embodiment, the pixel-level classification head is composed of a 3×3 convolutional layer and a 1×1 convolutional layer, and the convolutional layers all include batch normalization, and the activation function is RELU.

[0100] The specific process is as follows: The input image starts with a backbone block composed of two 3×3 convolutional layers, and then passes through 4 consecutive basic residual blocks to obtain a high-resolution feature map N, which can extract sufficient local information required for high-resolution feature maps for preprocessing. Then, the high-resolution feature map N and the low-resolution feature map n are respectively fed into the depth double-resolution module to achieve feature exchange between the high-resolution branch and the low-resolution branch, and the semantic representation of the high-resolution features is enhanced by means of the output of the low-resolution branch. The required low-resolution feature map n is obtained by dimensionality reduction of the high-resolution feature map N through a strided convolution. At this time, the stride of the low-resolution feature branch is 8, while the stride of the high-resolution feature branch always remains 8.

[0101] The depth double-resolution module is introduced in detail below:

[0102] The depth double-resolution module is as Figure 2 shown, and it includes a high-resolution branch, a low-resolution branch, and a bilateral fusion unit. The high-resolution branch includes two 3×3 convolutional operations and two batch normalization operations, and uses ReLU as the activation function. The high-resolution feature map N sequentially passes through a 3×3 convolutional operation, a batch normalization operation, a ReLU function, a 3×3 convolutional operation, and a batch normalization operation to extract feature information. At the same time, the high-resolution feature map N is connected with the output result of the last batch normalization operation by a residual connection, and the feature after the residual connection is denoted as N1; the number of channels of both 3×3 convolutional operations is 128.

[0103] The low-resolution branch has the same structure as the high-resolution branch, and also includes two 3×3 convolutional operations and two batch normalizations, and uses ReLU as the activation function. Only the number of channels of the two 3×3 convolutional operations is changed to 256, and the residual connection is also used. The input of the low-resolution branch is the low-resolution feature map n, and the feature after the residual connection is denoted as n1;

[0104] The bilateral fusion unit includes fusing the high-resolution branch into the low-resolution branch (high-to-low fusion) and fusing the low-resolution branch into the high-resolution branch (low-to-high fusion). For high-to-low fusion, N1 is downsampled through a 3×3 convolutional sequence (stride of 2 and number of channels of 256), and then after batch normalization, it is concatenated and fused with n1 to obtain n2.

[0105] For low-to-high fusion, n1 is first compressed through a 1×1 convolution, and after batch normalization, it is upsampled through bilinear interpolation and then concatenated and fused with N1 to obtain N2;

[0106] n2 uses the ReLU activation function, and after upsampling, it is fused with N2 processed by the ReLU activation function again to obtain a new high-resolution feature map F;

[0107] Meanwhile, n2 is processed using the ReLU activation function and undergoes stride convolution for dimensionality reduction to obtain a new low-resolution feature map f. At this time, the stride of the stride convolution is 16.

[0108] The FPformer module is introduced in detail below:

[0109] The FPformer module is as Figure 3 shown, and it is also a dual-resolution branch.

[0110] In the low-resolution branch, the new low-resolution feature map f passes through a GPU-friendly attention module to capture high-level global context. The GPU-friendly attention module can be expressed as:

[0111]

[0112] where Δ = 0.05 ∈ (-0.1, 0.1) is the bias correction coefficient used to adjust the GDN output, and α = 0.5 is the weight parameter used to adjust the influence of V. g GFA represents the GPU-friendly attention operation; GDN is the grouped double normalization operation; X represents the input feature, is the learnable parameter in the GPU-friendly attention module, M g = M × H, M is the parameter dimension, d is the feature dimension, H is the number of heads in the multi-head attention mechanism, and T represents matrix transpose.

[0113] After the attention map output by the GPU-friendly attention module is feature-fused with the new low-resolution feature map f, the fused feature map enters a KAN module and then undergoes a 3×3 convolution operation to extract features, and then is residually connected to the input of the KAN module to obtain the feature map X l .

[0114] In the high-resolution branch, a cross-attention module is used. The new high-resolution feature map F is concatenated and fused with the feature map X l to obtain the feature map X h , which is used as the input to the Q branch of the cross-attention module. At the same time, the feature map X l undergoes convolution and pooling operations to obtain the feature map X c , which is used as the input to the K and V branches of the cross-attention module. After being processed by the cross-attention module, the output attention map is residually connected to the input of the Q branch, and then passes through two stacked 3×3 convolutional layers to extract features. At the same time, batch normalization is used, the activation function uses the ReLU function, and residual connection is also used to obtain the feature map F1;

[0115] The high-level global context learned from the low-resolution branch is broadcast to each high-resolution pixel, making full use of the high-level semantic information learned from the low-resolution branch, and feeding the more representative features in the low-resolution branch into the cross-attention module in a stepped layout. The cross-attention module can be expressed as:

[0116]

[0117] where β = 1 is the weight parameter for adjusting the influence of V c . CA is the cross-attention module operation; X h , X l represent the feature maps on the high-resolution branch and the low-resolution branch respectively, is a set of matrix operations including splitting, permuting, and reshaping. After processing the feature map X c , the inputs to the K and V branches are denoted as K c , V c ; d h represents the feature dimension of the high-resolution branch. At the same time, the feature map X c is obtained by calculating X l through a function θ composed of a pooling layer and a convolutional layer. The spatial size of X c represents the number of tokens generated from X l .

[0118] The feature map F1 is concatenated with the feature map X l to obtain the low-resolution feature o as the output of the low-resolution branch; the low-resolution feature o is then upsampled and fused with the feature map F1 to generate the high-resolution feature O as the output of the high-resolution branch; at this time, the stride of the low-resolution feature branch is 32.

[0119] The following details the deep aggregation pyramid pooling module:

[0120] The deep aggregation pyramid pooling module is as shown in Figure 4 , including a pooling operation with a stride of 2 and a kernel of 5, a pooling operation with a stride of 4 and a kernel of 9, a pooling operation with a stride of 8 and a kernel of 17, and a global average pooling operation (with a kernel of H×W, where H is the height of the input image and W is the width of the input image);

[0121] The outputs after three pooling operations and global average pooling operation are each processed through a 1×1 convolution and an upsampling operation, resulting in a total of four branches; at the same time, the low-resolution feature o is processed through a 1×1 convolution to obtain the output y1. The output of y1 and the first branch is processed through a dilated convolution with a dilation rate of 4 to obtain the output y2. The output of y2 and the second branch is processed through a dilated convolution with a dilation rate of 8 to obtain the output y3. The output of y3 and the third branch is processed through a dilated convolution with a dilation rate of 12 to obtain the output y4. The output of y4 and the fourth branch is processed through a dilated convolution with a dilation rate of 16 to obtain the output y5; after y1, y2, y3, y4, and y5 are concatenated, they are processed through a 1×1 convolution and then concatenated with y1 to obtain the output of the deep aggregation pyramid pooling module.

[0122] After the low-resolution feature o is input, large pooling kernels with exponential strides are used to generate feature maps with image resolutions of 1 / 128, 1 / 256, and 1 / 512. And the input feature map generated by global average pooling and image-level information are utilized. Here, it is not sufficient to use a single 3×3 or 1×1 convolution to mix all multi-scale context information. First, the feature map is upsampled, and then more 3×3 dilated convolutions are used. Dilated convolutions with different dilation rates are used to receive receptive fields of different degrees, aiming to capture multi-scale context information, and at the same time, the context information of different scales is fused in a hierarchical residual manner. In addition, 1×1 projection shortcuts are added for easy optimization. In this module, the context extracted by the larger pooling kernel is integrated with the deeper information flow. By integrating different depths with pooling kernels of different sizes, a multi-scale property is formed. Finally, 1×1 convolutions are used to connect and compress all feature maps.

[0123] VI. Training and Calculation

[0124] 7000, 2000, and 1000 finely annotated images in the target dataset R are used for training, validation, and testing respectively. It is trained using the Adam W optimizer with an initial learning rate of 0.0004 and a weight decay of 0.0125. A poly learning strategy with an exponent of 0.9 is adopted to reduce the learning rate. The model is trained using 484 epochs (about 120K iterations), a batch size of 12, and syncBN batch normalization on 4 V100 GPUs. When the training model reaches convergence, that is, when the training gradient of the model is close to 0.

[0125] Measurement of over-excavation and under-excavation areas: The predicted probability map output by the model is binarized, with a threshold set to 0.5. Pixels with a probability greater than 0.5 are marked as over-excavated areas, and vice versa for under-excavated areas.

[0126] The area of the over-excavated or under-excavated area can be obtained by calculating the number of pixels within it:

[0127] S = N x Ap

[0128] N is the number of pixels in the area, A p is the actual area corresponding to each pixel. If the image resolution is R pixels / meter (i.e., there are R pixels per meter), then

[0129]

[0130] And for each connected area i, the area is:

[0131] S i = N i xA p

[0132] The total area is:

[0133]

[0134] where k is the number of areas.

[0135] Pre-input the labeled tunnel over-excavation and under-excavation image to be detected into the FP-Former semantic segmentation deep neural network for quantitative analysis to obtain the predicted probability map of the over-excavation and under-excavation areas, and then obtain the areas of the over-excavation and under-excavation areas. Calculate the ratio of the over-excavation or under-excavation area to the total area of the tunnel to get the area percentage, and set the area percentage threshold. When the detected area percentage of the over-excavation or under-excavation area exceeds the set area percentage threshold, the system immediately triggers an alarm. Set different levels of alarms according to the severity of the over-excavation or under-excavation area, such as:

[0136] ① Warning level: Slight over-excavation or under-excavation, the excavation posture of the excavator needs to be corrected.

[0137] ② Alarm level: Moderate over-excavation or moderate under-excavation, immediate measures need to be taken. Re-dig the under-excavated part and adjust the excavation posture.

[0138] ③ Emergency level: Severe over-excavation or severe under-excavation, the operation needs to be stopped and in-depth investigation carried out. Re-set the excavation plan and correct the deviation.

[0139] If it does not exceed the threshold range, the tunnel quality is qualified; otherwise, the quality is unqualified and an alarm is given and feedback is provided.

[0140] Specifically, check the area percentage. If the overexcavation area percentage is within [3%, 5%] of the total area or the underexcavation area percentage is within [2% - 3%] of the total area, set it as the warning level; if the overexcavation area percentage is within (5% - 10%] of the total area or the underexcavation area percentage is within (3% - 5%] of the total area, set it as the alarm level; if the overexcavation area percentage exceeds 10% of the total area or the underexcavation area percentage exceeds 5% of the total area, set it as the emergency level; if the remaining area ratio is not within the above ranges, consider it qualified.

[0141] The present invention makes a fair comparison between the FP-Former semantic segmentation deep neural network and other advanced methods, compares the segmentation effects of the tunnel overexcavation and underexcavation areas, evaluates on the above-mentioned constructed dataset (containing noisy images), uses mIoU and FPS as the measurement indicators for performance and efficiency respectively, and measures the FPS on a single RTX 2080Ti by default without tensor acceleration. The results of the experiment on the test set are shown in Table 1 below:

[0142] Table 1

[0143]

[0144] The present invention uses the FP-Former semantic segmentation deep neural network as the segmentation model for the target image, greatly improving the accuracy of recognition and quantification, enabling the system to achieve real-time high-accuracy quantification, with an FPS above 190 and an mIoU above 80%. The multi-scale dilated convolution combined with hierarchical residuals receives different feature information from multiple perspectives and performs feature fusion by combining context information, ensuring the quality of feature information while also solving the problems of information loss and the robustness of the model to noise in complex scenarios, making the model more stable and reliable. And by introducing the KAN module, the model can effectively reduce the risk of overfitting, while greatly reducing the amount of calculation and the number of parameters, and at the same time improving the efficiency and speed of the model. It can significantly reduce the size and computational requirements of the model while maintaining relatively high performance, especially in the industrial field environment, meeting the requirements of high real-time performance while maintaining high-precision segmentation performance.

[0145] As an embodiment, the present invention provides a real-time semantic segmentation tunnel overexcavation and underexcavation monitoring system, which is specifically used for detecting tunnel overexcavation and underexcavation and completing the quantitative analysis system, and specifically includes the following modules:

[0146] Data collection module: Use drones, laser scanners or high-definition camera devices to obtain high-resolution images during tunnel construction, ensure coverage of different excavation surface areas, and combine devices such as soil sensors and pressure sensors to record the real-time excavation depth and soil properties;

[0147] Data annotation module: Annotate the collected images, clearly distinguish the over-excavation and under-excavation image areas, and make accurate labels;

[0148] Data cleaning and enhancement module: Remove blurred, duplicate or invalid images to ensure the quality of the dataset. At the same time, increase the sample diversity through techniques such as rotation, scaling, cropping and color adjustment to improve the robustness of the model;

[0149] Image segmentation module: Establish an FP-Former semantic segmentation deep neural network for image segmentation, connect it with the data annotation module and the data cleaning and enhancement module to obtain the processed images, and use these to train the FP-Former semantic segmentation deep neural network for detecting the prediction probability map of the quantified tunnel over-excavation and under-excavation areas;

[0150] Over-excavation and under-excavation area measurement module: Binarize the prediction probability map, set the threshold to 0.5, mark the pixels with a probability greater than 0.5 as the over-excavation area, and vice versa as the under-excavation area, calculate the over-excavation and under-excavation area, calculate the ratio of the over-excavation or under-excavation area to the total tunnel area to obtain the area percentage, and set the area percentage threshold;

[0151] Early warning feedback module: Connect with the over-excavation and under-excavation area measurement module, and according to the detected area percentage, timely give an early warning to the background system and feedback the degree of tunnel over-excavation and under-excavation. The background system automatically issues an alarm and prompts relevant personnel to take measures.

[0152] The real-time tunnel over-excavation and under-excavation area monitoring system of the present invention has a high degree of customization, can be used in a variety of different working environments, and provides real-time feedback to the staff for timely adjustment of safety measures. It has wide applicability and practicability and can be used for safety monitoring in a variety of working environments, including tunnels, subway light rails, road bridges, and underground pipelines. In practical applications, the present invention provides a new idea and solution for problems such as the detection and real-time feedback of the earthwork over-excavation and under-excavation areas in a variety of workplaces, further promotes the development of this field, has great significance and broad application prospects, and also provides strong support and guarantee for the development of tunnel over-excavation and under-excavation monitoring technology.

[0153] Matters not described in the present invention are applicable to the prior art.

Claims

1. A real-time semantic segmentation method for monitoring overbreak and underbreak of tunnels, characterized in that The monitoring method includes the following steps: Step 1: Obtain high-resolution images of different excavation surface areas during tunnel construction suitable for excavation under different excavation conditions and geological backgrounds, and label the over-excavation and under-excavation area tags; Step 2: Data cleaning and enhancement After data cleaning and enhancement processing, a number of tunnel over-excavation and under-excavation images with a resolution of 512×1024 are obtained to form a target data set; Step 3: Construct an FP-Former semantic segmentation deep neural network The FP-Former semantic segmentation deep neural network includes a backbone block, a residual block, a deep double-resolution module, two stacked FPformer modules, a deep aggregation pyramid pooling module, and a pixel-level classification head; the backbone block is composed of two cascaded 3×3 convolutional layers, and each 3×3 convolutional layer includes a 3×3 convolution, a batch normalization operation, and a RELU operation; the residual block is composed of four consecutive basic residual blocks in series, the input of the residual block is the output of the backbone block, and the output of the residual block is a high-resolution feature map N. The high-resolution feature map N is downsampled by a stride convolution to obtain a low-resolution feature map n; the deep double-resolution module includes a high-resolution branch and a low-resolution branch, and the high-resolution feature map N and the low-resolution feature map n enter the high-resolution branch and the low-resolution branch of the deep double-resolution module respectively; the deep double-resolution module outputs a low-resolution feature f and a high-resolution feature F; then it enters two stacked FPformer modules, which output a high-resolution feature O and a low-resolution feature o; The low-resolution feature o enters the deep aggregation pyramid pooling module to extract context information, and the output feature information after processing by the deep aggregation pyramid pooling module is fused with the high-resolution feature O to obtain an output feature map with a stride of 8; Finally, the output feature map with a stride of 8 is passed to the pixel-level classification head for predicting dense semantic labels and outputting a prediction probability map; Step 4: Use the images in the target data set to train the FP-Former semantic segmentation deep neural network, use the trained FP-Former semantic segmentation deep neural network to monitor the over-excavation and under-excavation areas, and output a prediction probability map; perform binary processing on the output prediction probability map, set the threshold to 0.5, mark the pixels with a probability greater than 0.5 as over-excavation areas, and vice versa as under-excavation areas, so as to obtain the areas of the over-excavation or under-excavation areas respectively for engineering monitoring.

2. The monitoring method according to claim 1, wherein The pixel-level classification head is composed of a 3×3 convolutional layer and a 1×1 convolutional layer, and both convolutional layers contain batch normalization and ReLU activation functions.

3. The monitoring method according to claim 1, characterized in that, The deep double-resolution module includes a high-resolution branch, a low-resolution branch, and a bilateral fusion unit. The high-resolution branch includes two 3×3 convolution operations and two batch normalization operations, and uses ReLU as the activation function; the high-resolution feature map N sequentially passes through a 3×3 convolution operation, a batch normalization operation, a ReLU function, a 3×3 convolution operation, and a batch normalization operation to extract feature information. At the same time, the high-resolution feature map N is residually connected to the output result of the last batch normalization operation, and the feature after the residual connection is denoted as N1; The number of channels for both 3×3 convolution operations is 128; The low-resolution branch has the same structure as the high-resolution branch, also including two 3X3 convolution operations, two batch normalizations, using ReLU as the activation function and residual connections. However, the number of channels for the two 3×3 convolution operations is 256, and residual connections are also used. The input to the low-resolution branch is the low-resolution feature map n, and the feature after residual connection is denoted as n1; The bilateral fusion unit includes high-to-low fusion and low-to-high fusion. For high-to-low fusion, N1 is downsampled through a 3×3 convolution sequence with a stride of 2 and 256 channels, then batch-normalized and concatenated with n1 for a fusion operation to obtain n2; For low-to-high fusion, n1 is first compressed through a 1×1 convolution with 128 channels, batch-normalized, upsampled through bilinear interpolation, and then concatenated with N1 for a fusion operation to obtain N2; n2 is fused again with N2 processed by the ReLU activation function after passing through the ReLU activation function and upsampling to obtain a new high-resolution feature map F; At the same time, n2 is processed by the ReLU activation function and undergoes a strided convolution for dimensionality reduction to obtain a new low-resolution feature map f. At this time, the stride of the strided convolution is 16.

4. The monitoring method according to claim 1, characterized in that, The FPformer module has a low-resolution branch and a high-resolution branch. In the low-resolution branch, the new low-resolution feature map f passes through a GPU-friendly attention module to capture high-level global context. After the attention map output by the GPU-friendly attention module is feature-fused with the new low-resolution feature map f, the fused feature map enters a KAN module and then undergoes a 3×3 convolution operation to extract features, and then a residual connection is made with the input of the KAN module to obtain the feature map X l ; In the high-resolution branch, a cross-attention module is used. The new high-resolution feature map F and the feature map X l are concatenated and fused to obtain the feature map X h . The feature map X h is used as the input of the Q branch to the cross-attention module. At the same time, the feature map X l is processed by convolution and pooling to obtain the feature map X c . The feature map X c is used as the input of the K and V branches to the cross-attention module. After being processed by the cross-attention module, the output attention map is obtained. This attention map is then subjected to a residual connection with the feature map X h . After that, features are extracted through two stacked 3×3 convolutional layers, and a residual connection is made to the input of the two stacked 3×3 convolutional layers to obtain the feature map F1; Feature map F1 and feature map X l After splicing with the feature map X, a low-resolution feature o is obtained as the output of the low-resolution branch; the low-resolution feature o is then upsampled and fused with the feature map F1 to generate a high-resolution feature O as the output of the high-resolution branch.

5. The monitoring method according to claim 4, characterized in that The GPU-friendly attention module is expressed as: Among them, Δ = 0.05 ∈ (-0.1, 0.1) is the deviation correction coefficient for adjusting the GDN output; α = 0.5 is the weight parameter for adjusting the influence of V g ; GFA represents the GPU-friendly attention operation; GDN is the grouped double normalization operation; X represents the input feature; is the learnable parameter in the GPU-friendly attention module; M g = M × H, M is the parameter dimension, d is the feature dimension, H is the number of heads of the multi-head attention mechanism, and T represents the matrix transpose; The cross-attention module is expressed as: X c = θ(X l ) Among them, β = 1 is the weight parameter used to adjust the influence of V c ; CA is the cross-attention module operation; X h , X l respectively represent the feature maps on the high-resolution branch and the low-resolution branch. is a set of matrix operations including splitting, permuting, and reshaping, and the inputs to the K and V branches obtained after processing the feature map X c are respectively denoted as K c , V c ; d h represents the feature dimension of the high-resolution branch. At the same time, the feature map X c is obtained by calculating X l through the function θ composed of a pooling layer and a convolutional layer; the spatial size of X c represents the number of tokens generated from X l .

6. The monitoring method according to claim 1, wherein The deep aggregation pyramid pooling module includes pooling operations with a stride of 2 and a kernel of 5, pooling operations with a stride of 4 and a kernel of 9, pooling operations with a stride of 8 and a kernel of 17, and global average pooling operations; The outputs after the three pooling operations and the global average pooling operation are each processed through a 1×1 convolution and an upsampling operation, for a total of four branches; at the same time, the low-resolution feature o is processed through a 1×1 convolution to obtain the output y1. y1 and the output of the first branch are processed through a dilated convolution with a dilation rate of 4 to obtain the output y2. y2 and the output of the second branch are processed through a dilated convolution with a dilation rate of 8 to obtain the output y3. y3 and the output of the third branch are processed through a dilated convolution with a dilation rate of 12 to obtain the output y4. y4 and the output of the fourth branch are processed through a dilated convolution with a dilation rate of 16 to obtain the output y5; after y1, y2, y3, y4, and y5 are concatenated, they are processed through a 1×1 convolution and then concatenated with y1 to obtain the output of the deep aggregation pyramid pooling module.

7. The monitoring method according to claim 1, characterized in that, The FP-Former semantic segmentation deep neural network is trained using the Adam W optimizer. During training, the initial learning rate is 0.0004 and the weight decay is 0.0125; a poly learning strategy with a power of 0.9 is used to reduce the learning rate, and 484 epochs and a batch size of 12 are set.

8. The monitoring method according to claim 1, wherein The ratio of the overexcavation or under-excavation area to the total tunnel area is calculated to obtain the area percentage, and an area percentage threshold is set. When the detected overexcavation or under-excavation area percentage exceeds the set area percentage threshold, an alarm is immediately triggered, and different levels of alarms are set according to the severity of the overexcavation or under-excavation area. ① Warning level: Slight overexcavation or under-excavation, the excavator's digging posture needs to be corrected; ② Alarm level: Moderate overexcavation or moderate under-excavation. Immediate measures should be taken to re-excavate the under-excavated part and adjust the excavation posture. ③ Emergency level: Severe overexcavation or severe under-excavation. The operation needs to be stopped and in-depth investigation carried out, the excavation plan should be reset, and deviation correction should be made. If it does not exceed the area percentage threshold range, the tunnel quality is qualified; otherwise, the quality is unqualified and a warning is given and feedback is provided.

9. The monitoring method according to claim 7, wherein Check the area percentage. If the overexcavation area percentage is within [3%, 5%] of the total area or the under-excavation area percentage is within [2%-3%] of the total area, it is set as the warning level; if the overexcavation area percentage is within (5%-10%] of the total area or the under-excavation area percentage is within (3%-5%] of the total area, it is set as the alarm level; if the overexcavation area percentage exceeds 10% of the total area or the under-excavation area percentage exceeds 5% of the total area, it is set as the emergency level; if the remaining area ratio is not within the above range, it is regarded as qualified.

10. A real-time semantic segmentation tunnel overbreak and underbreak monitoring system, characterized in that, The system adopts the monitoring method described in claim 1 and includes the following modules: Data collection module: Use drones, laser scanners or high-definition camera devices to obtain high-resolution images during tunnel construction, ensure coverage of different excavation surface areas, and combine with soil sensors and pressure sensors to record the real-time excavation depth and soil properties. Data annotation module: Annotate the collected images to clearly distinguish the over / under-excavation image areas and make accurate labels. Data cleaning and enhancement module: Remove blurred, duplicate or invalid images, and at the same time perform image preprocessing using operations such as rotation, scaling, cropping, color adjustment, adding noise and flipping. Image segmentation module: Establish an FP-Former semantic segmentation deep neural network for image segmentation, connect it with the data annotation module and the data cleaning and enhancement module to obtain the processed images, and use these to train the FP-Former semantic segmentation deep neural network for detecting and quantifying the prediction probability map of the tunnel over / under-excavation area. Over / under-excavation area measurement module: Perform binary processing on the prediction probability map, set the threshold to 0.5, mark the pixels with a probability greater than 0.5 as the overexcavation area, and vice versa as the under-excavation area, calculate the over / under-excavation area, take the ratio of the over / under-excavation area to the total tunnel area to obtain the area percentage, and set the area percentage threshold. Warning and feedback module: Connect with the over / under-excavation area measurement module, and according to the detected area percentage, timely give a warning to the background system and feedback the degree of tunnel over / under-excavation. The background system automatically issues an alarm and prompts relevant personnel to take measures.

Citation Information

Patent Citations

  • Tunnel crack detection and measurement method based on dual-depth learning model

    CN112508030A

  • Method, system and device for measuring over-break and under-break in tunnel blasting and medium

    CN118823094A