Tunnel surrounding rock crack detection and direction prediction method based on improved YOLO model
By improving the YOLO model, combining dynamic upsampling and lightweight attention mechanism, and combining probabilistic Hough transform and geometry-aware bounding box loss function, the accuracy problems of crack detection and direction estimation are solved, and higher-precision crack detection and direction prediction are achieved.
Patent Information
- Application Number
- CN202510949810.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies have difficulty accurately detecting small or broken cracks in low-quality or noisy images, especially in environments such as underground tunnels and long-span bridges, and the crack direction estimation error is large, affecting the overall reliability of the structure.
An improved YOLO model is adopted, combined with a dynamic upsampling module, a lightweight attention mechanism, and multi-scale feature fusion. Probabilistic Hough transform is used for crack detection and direction prediction. A geometrically aware bounding box loss function is introduced to improve detection accuracy and direction estimation accuracy.
The accuracy of crack detection and the reliability of direction estimation are significantly improved, and the application potential of structural monitoring is enhanced, especially in complex and multi-scale crack environments, providing more accurate crack positioning and direction classification.
Smart Images

Figure CN120766039A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of non-destructive testing technology, and in particular to a method for detecting and predicting cracks in surrounding rock of a tunnel based on an improved YOLO model. Background Art
[0002] Crack detection and orientation estimation in RGB imagery plays a key role in infrastructure inspection, structural health monitoring, and maintenance management. Accurate and timely identification of cracks and their directional properties is crucial for assessing the structural integrity and safety of critical civil engineering assets such as bridges, tunnels, roads, and buildings. Early detection and precise crack characterization facilitate preventative maintenance, reduce the risk of structural failure, and extend the life of infrastructure.
[0003] Deep learning can be used in the field of crack detection by providing robust feature representations and enhanced generalization capabilities. Object detection frameworks such as YOLO, Faster R-CNN, and Mask R-CNN have shown promising results in automated crack detection. However, they are less effective in accurately detecting thin, small, or broken cracks in low-quality or noisy images, and under the real-time constraints faced in common environments such as underground tunnels and long-span bridges.
[0004] In addition to crack location, crack direction information is crucial for assessing stress distribution and potential failure modes in structures. The primary direction of cracks can help engineers predict crack propagation trends and prioritize repair measures, thereby improving structural safety.
[0005] However, many existing techniques treat direction estimation as a secondary task, often relying on simple geometric models or post-processing steps to infer the crack direction. These methods are prone to introducing errors, reducing the accuracy of direction estimation and thus affecting the overall reliability of the system. Summary of the Invention
[0006] In order to overcome the above problems existing in the prior art, the present invention proposes a tunnel surrounding rock crack detection and direction prediction method based on an improved YOLO model.
[0007] The technical solution adopted by the present invention to solve the technical problem is: a method for detecting and predicting the direction of tunnel surrounding rock cracks based on an improved YOLO model, comprising the following steps: Step 1: construct a crack detection model; Step 2: input the image to be tested into the crack detection model obtained in step 1, and the crack detection model identifies candidate crack areas; Step 3: extract the corresponding image blocks of the candidate crack areas identified in step 2, convert them into grayscale images, and apply an edge detector to extract edge features to obtain an edge map; Step 4: Extract geometric primitives from the edge graph through probabilistic Hough transform to generate a set of line segment candidates; Step 5: Select the longest line segment from the line segment candidate set obtained in step 4 to determine the main direction of the crack, and perform direction classification based on the main direction.
[0008] In the above-mentioned method for detecting and predicting cracks in surrounding rock of tunnels based on the improved YOLO model, the crack detection model in step 1 is an improved YOLOv11 model. The improvements of the improved YOLOv11 model include: replacing the standard upsampling layer of the neck with a dynamic upsampling layer, using a cascaded dynamic head as the detection head, and introducing a geometrically aware bounding box loss function.
[0009] In the above-mentioned tunnel surrounding rock crack detection and direction prediction method based on the improved YOLO model, the dynamic upsampling layer includes a kernel prediction module and a content-aware reconstruction module; the cascaded dynamic head is formed by stacking multiple feature-aware detection heads, and the feature-aware detection head introduces a scale-aware attention mechanism, a space-aware attention mechanism, and a task-aware attention mechanism.
[0010] In the above-mentioned tunnel surrounding rock crack detection and direction prediction method based on the improved YOLO model, the kernel prediction module projects the input feature map into a low-dimensional space, applies a convolution operation on the reduced-dimensional feature map, generates a position-specific convolution kernel vector for upsampling, and reshapes and normalizes the predicted convolution kernel vector at each spatial position. The content-aware reconstruction module extracts a local neighborhood from the input feature map with the source position corresponding to each position calculator of the output feature map in the input feature map as the center, applies the convolution kernel vector obtained by the kernel prediction module to the local neighborhood, and calculates the upsampled value of the output position.
[0011] In the above-mentioned tunnel surrounding rock crack detection and direction prediction method based on the improved YOLO model, the scale-aware attention mechanism performs average pooling in the spatial dimension and channel dimension to obtain a compact descriptor of each layer feature, inputs the descriptor into the convolution layer, performs nonlinear transformation, uses the HardSigmoid function to generate scale attention weights, and applies the generated scale attention weights to the original input features to obtain features. ; The spatially aware attention mechanism predicts the spatial offset for deformable convolution on the input feature map, applies a convolution operation to generate a spatial attention map, and applies the generated spatial attention map to the scale-weighted feature map. , and obtain the spatially enhanced features ; The task-aware attention mechanism extracts global context information by performing average pooling in the scale dimension and the spatial dimension, and inputs the pooled vector into two fully connected layers in sequence to finally obtain the channel attention vector, which is applied to the spatially enhanced features. , and obtain the final output features .
[0012] In the above-mentioned tunnel surrounding rock crack detection and direction prediction method based on the improved YOLO model, the geometric perception bounding box loss function is specifically: ; ; ; ; ; in, The center point distance term representing direction perception; Represents the penalty term for the target width and height shape deviation; the coefficient 0.5 is used to balance the impact of the shape deviation term on the overall loss. express Loss function; and Represent the center coordinates of the predicted box and the real box respectively; Represents the diagonal length of the minimum bounding box; weight term and is the direction perception factor; Indicates the width of the real frame, scale is the scaling factor, Indicates the height of the real frame, Indicates the width of the real frame after scaling, Indicates the height of the real frame after scaling, and represents the directional shape deviation term; w represents width, and h represents height.
[0013] In the above-mentioned tunnel surrounding rock crack detection and direction prediction method based on the improved YOLO model, step 4 is specifically: using probabilistic Hough transform to extract geometric primitives from the edge map and generate a set of line segment candidates: ; Each element represents a line from point Arrive Line segments, the direction of each line segment is calculated by its angle with the horizontal axis: ; In order to determine the main direction of the crack, the longest line is selected from all the line segments, and its direction angle is defined as the main direction : ; Where j represents the index of the geometric unit, and M represents the set of all geometric units in the edge graph.
[0014] The above-mentioned tunnel surrounding rock crack detection and direction prediction method based on the improved YOLO model, the step 5 is specifically as follows: The value of categorizes the crack into one of three orientations: .
[0015] The beneficial effect of the present invention is that it proposes a robust unified framework for simultaneous crack detection and direction estimation in RGB images. Based on the Crack-YOLOv11 architecture, the present invention combines a dynamic upsampling module, a lightweight attention mechanism, and multi-scale feature fusion to effectively improve the detection capability of small-sized and irregularly shaped cracks. At the same time, a geometric analysis module based on probabilistic Hough transform is introduced to accurately estimate the main direction of the crack based on detection. This end-to-end framework not only significantly improves detection accuracy, but also provides reliable direction estimation, enhancing its application potential in actual structural monitoring.
[0016] This paper replaces traditional upsampling techniques with a dynamic upsampling module, enabling the model to dynamically adjust to changes in local context, significantly improving the recovery of details in small crack regions. Furthermore, the feature-aware module combines scale-aware, spatial-aware, and task-aware attention mechanisms to adaptively assign importance weights to features at different levels, improving the ability to recognize complex and multi-scale cracks.
[0017] To achieve better tunnel crack detection, this paper introduces an enhanced positioning loss function based on CloU. This loss function combines angle alignment and shape similarity to effectively improve the accuracy of crack positioning.
[0018] The detected crack masks are post-processed through edge detection and line segment analysis to achieve accurate crack direction classification. This method provides important direction information for structural diagnosis and hazard warning. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic flow chart of the present invention; Figure 2 It is a structural diagram of the crack detection model of the present invention; Figure 3 Schematic diagram of crack detection results of different models in an embodiment of the present invention. DETAILED DESCRIPTION
[0020] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] This embodiment discloses a method for detecting and predicting cracks in surrounding rock of a tunnel based on an improved YOLO model. Figure 1 As shown, specifically including: Step 1: Build a crack detection model.
[0022] The crack detection model in this embodiment is an improved YOLOv11 model, and the improvements of the improved YOLOv11 model include: first, the standard upsampling layer in the neck is replaced by a dynamic upsampling module, which dynamically generates an upsampling kernel based on the local feature context. This modification significantly enhances the model's ability to detect small targets with higher accuracy. Secondly, the detection head is redesigned using a feature-aware detection head, which incorporates a multi-dimensional attention mechanism to adaptively strengthen feature aggregation and improve spatial semantic representation. Third, the loss function is optimized by integrating a geometric-aware bounding box loss metric, which takes into account the geometric characteristics of the predicted bounding box. This adjustment can better handle targets with extreme aspect ratios and improve the positioning accuracy of small targets. The architecture of the crack detection model in this embodiment is as follows Figure 2 shown.
[0023] The dynamic upsampling module generates content-sensitive upsampling kernels based on the local context of the feature map, resulting in higher quality upsampled outputs with higher semantic and spatial fidelity. The dynamic upsampling module consists of two key components: the kernel prediction module and the content-aware reconstruction module. Given an input feature map ,in and Represents height, width and number of channels respectively. The goal is to Upsampling scale factor To obtain the output feature map .
[0024] (1) Kernel prediction module To reduce computational cost, the input feature map is first projected into a low-dimensional space: ; in, Represents the number of intermediate channels after dimensionality reduction. This operation significantly reduces computational overhead while still retaining important contextual information. Apply a The convolution operation is used to generate the position-specific convolution kernel weights for upsampling: ; Among them, K is the convolution kernel size used for reconstruction. Indicates from a Reconstruct a The number of weights required for high-resolution tiles.
[0025] For each spatial position The convolution kernel vector predicted above , reshape it and normalize it using the Softmax function: .
[0026] This normalization operation ensures that the weights are non-negative and sum to 1, thus forming an effective probability weighting function to support the subsequent reconstruction process.
[0027] (2) Content-aware reconstruction module In the reconstruction phase, for the output feature map Each position in , first calculate its corresponding source position in the input feature map. , and its corresponding input feature map position is: ; With this location As the center, from the input feature map Extract one The local neighborhood of : ; in, Indicates Central The position set of all pixels in the neighborhood. The output position is calculated by applying the predicted convolution kernel to the local image block. Upsampled values on: ; The kernel weights are shared across channels to maintain computational efficiency and reduce the number of parameters.
[0028] The feature-aware detection head architecture reconstructs the input into a 3D tensor where represents the number of layers of the feature pyramid, represents the spatial resolution after flattening, is the number of feature channels. In order to effectively capture the dependencies between these three dimensions, the feature-aware detection head introduces three attention mechanisms: scale-aware attention , spatial perception and attention and task-aware attention These three modules work together in scale, space, and channel dimensions to improve the performance of object detection.
[0029] (1) Scale-aware attention mechanism The scale-aware attention mechanism adaptively weights features from different pyramid layers to better adapt to objects of different sizes. First, in the spatial dimension and channel dimension Average pooling is performed on the layers to obtain a compact descriptor of each layer feature. The descriptor is then input into a Convolutional layer, and nonlinear transformation through ReLU activation function. Finally, the HardSigmoid function is used to generate scale attention weights: The obtained attention weight is , these weights are then broadcast and applied to the original input features: ; in, Represents element-wise multiplication with appropriate broadcasting.
[0030] (2) Spatial Perception Attention Mechanism The spatially aware attention module enhances discriminative features by focusing on salient spatial regions and suppressing irrelevant background information. The module first predicts the spatial offset for deformable convolution on the input feature map X, and then applies a 3×3 convolution operation: .
[0031] This operation generates a spatial attention map This spatial attention map is then applied to the scale-weighted features , to obtain the spatially enhanced features: .
[0032] (3) Task-aware attention mechanism To balance the objectives between classification and regression tasks in object detection, the task-aware attention module focuses on optimizing feature representation in the channel dimension.
[0033] First, by using the scale dimension and spatial dimensions Average pooling is performed on the vector to extract global context information. Then, the pooled vector is input into two fully connected layers in sequence, with ReLU activation function used in the middle and normalization processing: The final channel attention vector is Finally, the channel attention is applied to the spatially enhanced features , and get the final output features: .
[0034] (4) Cascade dynamic header module The various attention mechanisms described above are integrated into a unified module called the Feature-Aware Detection Head. To enhance the expressive power of the model, multiple such modules are stacked.
[0035] Suppose the total stacked A module with the following recursive form: .
[0036] Finally passed The output after stacking modules is: .
[0037] This architecture enables the detection head to dynamically focus on the most informative scales, spatial regions, and task-related features at multiple stages, thereby effectively improving detection accuracy.
[0038] This example also introduces a geometry-aware bounding box loss function. This loss function expands on CIoU by adding a direction-aware center point distance penalty term and a shape deviation term to better align the predicted box with the irregular shape of the crack-like object. The proposed loss function is as follows: ; in, The center point distance term representing the direction perception; Represents the penalty term for the target width and height shape deviation; the coefficient 0.5 is used to balance the impact of the shape deviation term on the overall loss.
[0039] The first term quantifies the predicted box With real box The degree of overlap between: .
[0040] To better consider the directionality and scale differences, we introduce a shape-weighted center point distance term: ; in, and Represent the center coordinates of the predicted box and the real box respectively; Indicates the diagonal length of the minimum bounding box. Weight term and is a direction-aware factor that emphasizes the importance of directional alignment based on the shape of the target: ; To penalize the deviation in width and height, a shape deviation term is defined: ; Among them, the direction shape deviation term and They are defined as follows: ; By combining the geometric alignment term with the shape-sensitive constraint term, the geometry-aware bounding box loss function significantly improves the bounding box regression performance for slender or irregularly shaped targets, which is particularly suitable for application scenarios with irregular targets such as tunnel surrounding rock crack detection.
[0041] Step 2: Input the image to be tested into the crack detection model obtained in step 1, and the crack detection model identifies candidate crack areas.
[0042] Given an input image First, the real-time target detector in the previous stage is used to identify the candidate crack regions and obtain A collection of bounding boxes: ; in, Indicates the The center coordinates of the bounding box, and are width and height respectively, is the confidence score, which indicates the probability that the bounding box contains a crack. Subsequently, non-maximum suppression is applied to remove redundant detection boxes.
[0043] Step 3: extract the corresponding image blocks of the candidate crack areas identified in step 2, convert them into grayscale images, and apply edge detectors to extract edge features to obtain edge maps.
[0044] For each region of interest, the corresponding image block is extracted and converted into a grayscale image for boundary analysis. Then the Canny edge detector is applied to enhance edge features and suppress noise: ; in, It is the grayscale ROI image. is the generated binary edge map. Threshold and It is selected through empirical methods to strike a balance between detection sensitivity and robustness.
[0045] Step 4: Extract geometric primitives from the edge graph through probabilistic Hough transform to generate a set of line segment candidates.
[0046] Use the probabilistic Hough transform to extract geometric primitives from the edge map and generate a set of line segment candidates: ; Each element represents a line from point Arrive Line segments. The direction of each line segment is calculated by its angle with the horizontal axis, with the unit of angle being degrees: .
[0047] In order to determine the main direction of the crack, the longest line is selected from all the line segments, and its direction angle is defined as the main direction : .
[0048] Step 5: Select the longest line segment from the line segment candidate set obtained in step 4 to determine the main direction of the crack, and perform direction classification based on the main direction.
[0049] according to The values of can be roughly classified into one of three directions: .
[0050] Although this classification method is relatively rough, it can provide valuable assessment of crack directionality, which is crucial for analyzing structural stress distribution patterns and potential failure risks.
[0051] Based on the above method, this example conducted experimental verification. The specific experimental settings are as follows: the experiment was conducted on a workstation equipped with the Ubuntu 22.04 operating system and an NVIDIA RTX 3090 graphics card. This hardware configuration provides sufficient computing resources for deep learning tasks. The model was written in Python 3.12 and implemented in the PyTorch 2.3.0 framework, using CUDA 12.1 for GPU acceleration.
[0052] To verify the effectiveness of the proposed improvements, we conducted a comprehensive comparative study of the model in this embodiment (the Crack-YOLO model) with several classic object detection frameworks, including RetinaNet, YOLOv5, YOLOv8, RT-DETR, Faster R-CNN, EfficientDet-D3, DETR, YOLOX, and the baseline YOLOv11. The results are summarized in Table 1.
[0053] Table 1 Comparison of different target detection models
[0054] Table 1 demonstrates the comprehensive comparison of various target detection models in the crack detection task, evaluating precision, recall, mAP@50, and mAP@50-95. Crack-YOLO achieves the highest precision (88.1%) and recall (85.4%), significantly outperforming other models. Compared to RetinaNet, its precision improves by 10.0 percentage points (78.1% vs. 88.1%) and recall by 10.8 percentage points (74.6% vs. 85.4%), highlighting that Crack-YOLO can accurately identify cracks while effectively reducing false positives and false negatives.
[0055] Furthermore, Crack-YOLO also improves upon its backbone model YOLOv11, with mAP@50 increasing by 6.3 percentage points (79.8% vs. 86.1%) and mAP@50-95 by 6.6 percentage points (53.2% vs. 59.8%). This underscores the effectiveness of the model in handling various crack detection scenarios, particularly in terms of accurate localization and overall performance.
[0056] While YOLOX ranks second in precision (85.4%) and recall (77.8%), it still lags significantly behind Crack-YOLO in mAP@50-95 (56.2% vs. 59.8%). This indicates that although YOLOX performs well in detection accuracy, it struggles with precise localization of smaller or overlapping cracks, where Crack-YOLO excels.
[0057] RT-DETR and YOLOv11 also show competitiveness, with RT-DETR achieving 84.0% precision and 76.5% recall, but its mAP@50-95 values (51.1% and 53.2%, respectively) reveal shortcomings in precise crack localization. In contrast, traditional models like FasterR-CNN and DETR perform lower, with FasterR-CNN achieving 77.8% precision and 47.7% mAP@50-95, and DETR achieving 79.2% precision and 44.5% mAP@50-95.
[0058] Crack-YOLO emerges as the leading model for crack detection, consistently outperforming other advanced models in terms of detection precision, recall, and localization accuracy. Its significant improvement over existing models, particularly in mAP@50-95, demonstrates the robustness of Crack-YOLO in accurately detecting and localizing cracks under complex or overlapping conditions. Therefore, Crack-YOLO is the most reliable and efficient model for the crack detection task when compared to other contemporary models.
[0059] This example selects four models (RetinaNet, RT-DETR, YOLOv8, and Crack-YOLO) for crack detection tasks. The experimental results are as follows: Figure 3 As shown in the figure. In the first column, YOLOv8 fails to detect the crack, while RT-DETR and RetinaNet are able to detect the crack but show some inaccuracy, indicating some of their limitations in detection. The results in the second and third columns show that although all models detected the crack, Crack-YOLO achieved the highest probability value, indicating its superior performance in crack identification. Finally, in the fourth column, RT-DETR failed to detect the crack. Although YOLOv8 has a higher detection probability than Crack-YOLO, overall, Crack-YOLO outperforms other models in detection accuracy and stability. In summary, Crack-YOLO performs well in crack detection and has a superior overall performance.
[0060] In addition, this embodiment also evaluates the ability of the method disclosed in this embodiment to classify crack directions. The evaluation results are shown in Table 2.
[0061] Results demonstrate that the proposed method achieves robust and balanced classification across all orientation categories, with an overall accuracy of 88.4% and a macro-average F1 score of 88.3%. Precision and recall values demonstrate consistent performance across horizontal, vertical, and tilted cracks, indicating no significant bias toward any particular orientation. The high macro-average F1 score highlights the effectiveness of the orientation classifier, particularly when dealing with ambiguous and irregular crack patterns. This underscores the advantage of incorporating segment geometry rather than relying solely on global texture or shape features.
[0062] Table 2 Direction classification performance
[0063] In summary, this embodiment significantly enhances the resolution and detail recovery capabilities of the feature map by integrating an upsampling module based on a dynamic kernel, thereby improving the detection accuracy of small-scale cracks. The feature-aware detection head introduces a multi-dimensional attention mechanism, which effectively integrates multi-scale and spatial semantic features. In addition, this paper designs a geometry-aware bounding box loss function, which combines angle alignment and shape matching techniques to reduce the error in tunnel crack positioning. Combined with the edge and line segment extraction technology based on probabilistic Hough transform, this method achieves accurate positioning and direction classification of cracks. Experimental results show that the proposed method is significantly superior to traditional methods in crack detection accuracy and direction estimation, demonstrating its potential in the safety assessment of tunnel engineering structures.
[0064] The above embodiments are merely exemplary embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art may make various modifications or equivalent substitutions to the present invention within the spirit and scope of protection of the present invention, and such modifications or equivalent substitutions shall also be deemed to fall within the scope of protection of the present invention.
Claims
1. A tunnel surrounding rock crack detection and direction prediction method based on an improved YOLO model is characterized by: The steps include: Step 1: construct a crack detection model; Step 2: input the image to be tested into the crack detection model obtained in step 1, and the crack detection model identifies candidate crack areas; Step 3: extract the corresponding image blocks of the candidate crack areas identified in step 2, convert them into grayscale images, and apply an edge detector to extract edge features to obtain an edge map; Step 4: Extract geometric primitives from the edge graph through probabilistic Hough transform to generate a set of line segment candidates; Step 5: Select the longest line segment from the line segment candidate set obtained in step 4 to determine the main direction of the crack, and perform direction classification based on the main direction.
2. The method for detecting and predicting cracks in surrounding rock of a tunnel based on the improved YOLO model according to claim 1, characterized in that: The crack detection model in step 1 is an improved YOLOv11 model. The improvements of the improved YOLOv11 model include: replacing the standard upsampling layer of the neck with a dynamic upsampling layer, using a cascaded dynamic head as the detection head, and introducing a geometrically aware bounding box loss function.
3. The method for detecting and predicting cracks in surrounding rock of a tunnel based on the improved YOLO model according to claim 2, characterized in that: The dynamic upsampling layer includes a kernel prediction module and a content-aware reconstruction module; the cascaded dynamic head is formed by stacking multiple feature-aware detection heads, and the feature-aware detection head introduces a scale-aware attention mechanism, a space-aware attention mechanism, and a task-aware attention mechanism.
4. The method for detecting and predicting cracks in surrounding rock of a tunnel based on the improved YOLO model according to claim 3, characterized in that: The kernel prediction module projects the input feature map into a low-dimensional space and applies a convolution operation on the reduced-dimensional feature map to generate a position-specific convolution kernel vector for upsampling. The predicted convolution kernel vector at each spatial position is reshaped and normalized. The content-aware reconstruction module extracts a local neighborhood from the input feature map with the source position corresponding to each position calculator of the output feature map in the input feature map as the center, applies the convolution kernel vector obtained by the kernel prediction module to the local neighborhood, and calculates the upsampled value of the output position.
5. The method for detecting and predicting cracks in surrounding rock of a tunnel based on the improved YOLO model according to claim 3, characterized in that: The scale-aware attention mechanism performs average pooling in the spatial dimension and channel dimension to obtain a compact descriptor of each layer feature, inputs the descriptor into the convolution layer, performs nonlinear transformation, uses the HardSigmoid function to generate scale attention weights, and applies the generated scale attention weights to the original input features to obtain features ; The spatially aware attention mechanism predicts the spatial offset for deformable convolution on the input feature map, applies a convolution operation to generate a spatial attention map, and applies the generated spatial attention map to the scale-weighted feature map. , and obtain the spatially enhanced features ; The task-aware attention mechanism extracts global context information by performing average pooling in the scale dimension and the spatial dimension, and inputs the pooled vector into two fully connected layers in sequence to finally obtain the channel attention vector, which is applied to the spatially enhanced features. , and obtain the final output features .
6. The method for detecting and predicting cracks in surrounding rock of a tunnel based on the improved YOLO model according to claim 2, characterized in that: The geometry-aware bounding box loss function is specifically: ; ; ; ; ; in, The center point distance term representing the direction perception; Represents the penalty term for the target width and height shape deviation; the coefficient 0.5 is used to balance the impact of the shape deviation term on the overall loss. express Loss function; and Represent the center coordinates of the predicted box and the real box respectively; Represents the diagonal length of the minimum bounding box; weight term and is the direction perception factor; Indicates the width of the real box, Indicates the height of the real frame, scale is the scaling factor, Indicates the width of the real frame after scaling, Indicates the height of the real frame after scaling, and represents the directional shape deviation term; w represents width, and h represents height.
7. The method for detecting and predicting cracks in surrounding rock of a tunnel based on the improved YOLO model according to claim 1, characterized in that: The step 4 is specifically: using probabilistic Hough transform to extract geometric primitives from the edge map and generate a set of line segment candidates: ; Each element represents a line from point Arrive Line segments, the direction of each line segment is calculated by its angle with the horizontal axis: ; In order to determine the main direction of the crack, the longest line is selected from all the line segments, and its direction angle is defined as the main direction : ; Where j represents the index of the geometric unit, and M represents the set of all geometric units in the edge graph.
8. The method for detecting and predicting cracks in surrounding rock of a tunnel based on the improved YOLO model according to claim 1, characterized in that: The step 5 is specifically as follows: The value of categorizes the crack into one of three orientations: 。