Tunnel anomaly detection model establishment and detection early warning method
By improving the backbone network and adaptive rectangular convolution module of the YOLOv10 model, combined with the SADMA module and WGIoU loss function, the problems of low accuracy and high computational complexity of traditional tunnel detection methods in low light and high temperature environments are solved, and efficient tunnel anomaly detection on edge computing devices is realized.
Patent Information
- Application Number
- CN202510876675.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-27
AI Technical Summary
Traditional tunnel detection methods have low detection accuracy in low light or high temperature environments, and the three-dimensional lidar calculations are large, making it difficult to operate efficiently on edge computing devices.
Using the ARDMADet model, the backbone network and adaptive rectangular convolution module of the YOLOv10 model are improved, combined with the SADMA module and WGIoU loss function, efficient processing and detection of three-dimensional point cloud data is achieved.
It significantly reduces the computational complexity, improves the accuracy and robustness of small target detection, adapts to different scales and shapes of targets, and is suitable for resource-constrained edge computing platforms.
Smart Images

Figure CN120388294A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of tunnel anomaly detection, and specifically to the establishment of a tunnel anomaly detection model and a detection and early warning method. Background Art
[0002] Performing personnel detection and alarm in a tunnel environment is a key safety technology, mainly applied in fields such as rail transit, mine tunnels, and underground engineering. Traditional monitoring means mainly include video surveillance systems, infrared sensors, and ultrasonic detection devices. Video surveillance systems are difficult to provide clear images in low-light or strong dust environments, and are susceptible to light changes, resulting in a decrease in detection accuracy. Infrared sensors may not be able to accurately detect personnel in high-temperature environments or in the presence of heat source interference, while ultrasonic sensors may produce false positives under the influence of tunnel echoes.
[0003] In addition, although three-dimensional lidar (LiDAR), as a high-precision detection device, can provide complete spatial information, traditional three-dimensional point cloud data processing has a large amount of calculation and high hardware resource requirements, making it difficult to operate efficiently on edge computing devices. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for establishing a tunnel anomaly detection model and a detection and early warning method that reduce computational complexity.
[0005] To achieve the above object, the present invention adopts the following technical solution: A method for establishing a tunnel anomaly detection model, establishing an ARDMADet model, which is an improved YOLOv10 model. The ARDMADet model includes a backbone network for feature extraction of the input image, a neck network for feature extraction and fusion of the feature map, and a head structure for detecting and classifying the fused feature map output by the neck network; Replace the C2f module in the backbone network of the YOLOv10 model with a feature extraction unit SADMA_AR module, which is improved based on the C2f module; The feature extraction unit SADMA_AR module adds a first SADMA module in the Concat module and the convolutional module of the C2f module, replaces both convolutional modules in the C2f module with a first adaptive rectangular convolutional module, the feature map output by the first adaptive rectangular convolutional module is sequentially subjected to normalization processing and first activation function processing, and replaces the Bottleneck module in the C2f module with a Mablock module; The Mablock module includes two second adaptive rectangular convolution modules, a second SADMA module, and a connection module that are connected in sequence. The feature maps output by the second adaptive rectangular convolution module are sequentially processed by normalization and a first activation function. The connection module is used to perform connection processing on the feature map input to the Mablock module and the feature map output by the second SADMA module.
[0006] Preferably, both the first SADMA module and the second SADMA module include four branches: The first branch rotates the initial feature map by the first rotation to obtain a first rotated feature map; The second branch performs max-pooling and average-pooling operations on the first rotated feature map, uses deformable convolution modules with different dimensions to extract features from the pooled feature maps respectively, splices the feature maps with different dimensions, and outputs them using a second activation function; Perform an element-wise multiplication operation on the first rotated feature map and the feature map output by the second branch, and then perform the first reverse rotation and output; The third branch rotates the initial feature map by the second rotation to obtain a second rotated feature map, performs max-pooling and average-pooling operations on the second rotated feature map, uses deformable convolution modules with different dimensions to extract features from the pooled feature maps respectively, splices the feature maps with different dimensions, and outputs them using a second activation function, and then performs the second reverse rotation on the feature map output by the second activation function and outputs; The fourth branch performs max-pooling and average-pooling operations on the initial feature map, uses deformable convolution modules with different dimensions to extract features from the pooled feature maps respectively, splices the feature maps with different dimensions, and outputs them using a second activation function, and performs an element-wise multiplication operation on the feature map output by the second activation function and the initial feature map and then outputs; Add the feature map output after the first reverse rotation, the feature map output after the second reverse rotation, and the feature map output by the fourth branch, and then calculate the average value and output.
[0007] Preferably, the first rotation is a 90° counterclockwise rotation along the H axis, the first reverse rotation is a 90° clockwise rotation along the H axis, the second rotation is a 90° counterclockwise rotation along the W axis, and the second reverse rotation is a 90° clockwise rotation along the W axis.
[0008] Preferably, the second activation function is a Sigmoid activation function.
[0009] Preferably, the first activation function is a SILU activation function.
[0010] The tunnel anomaly detection and warning method includes the following steps that are executed sequentially: S1: Obtain the three-dimensional point cloud data in the tunnel in real time. The three-dimensional point cloud data is converted through projection to obtain a two-dimensional data set, and the two-dimensional data set is preprocessed, and the preprocessed two-dimensional data set is divided into a training set, a validation set and a test set according to a preset ratio; S2: Input the training set into the ARDMADet model established by using the tunnel anomaly detection model establishment method described in any one of the above to train the ARDMADet model; S3: Use the WGIoU loss function to perform backward adjustment on the trained ARDMADet model to obtain the adjusted ARDMADet model. The calculation formula of the WGIoU loss function is as follows: ; Among them, A represents the ground truth box, B represents the preset bounding box, C represents the smallest box containing the ground truth box and the predicted box, represents the coordinates of the center of the anchor box, represents the coordinates of the center of the target box, and represent the size of the minimum bounding box, , where , is the outlier degree, is the IoU threshold between the predicted bounding box and the ground truth bounding box, is 's moving average value. A small outlier degree means a high-quality anchor box. and are hyperparameters, is constructed a non-monotonic focusing coefficient; S4: Input the test set into the adjusted ARDMADet model for testing to obtain the test results; S5: If the test results show that there are people in the preset area, an alarm prompt is issued.
[0011] Preferably, the preprocessing in step S1 includes sequentially performing background filtering, ROI region cropping and noise suppression on the two-dimensional data set.
[0012] A computer product includes a computer program, and when the computer program is executed by a processor, it implements the tunnel anomaly detection and warning method described in any one of the above.
[0013] A computer storage medium stores a computer program, and when the computer program is executed by a processor, it implements the tunnel anomaly detection and warning method described in any one of the above.
[0014] By adopting the foregoing design scheme, the beneficial effects of the present invention are as follows: The ARDMADet model established in this application embeds the adaptive rectangular convolution module into the C2f module. The adaptive rectangular convolution module can dynamically learn the height and width of the convolution kernel to generate a rectangular convolution kernel, thereby flexibly adjusting the shape of the convolution window according to the sizes of different objects in the image. Moreover, the adaptive rectangular convolution module can not only change the shape of the convolution kernel, but also dynamically adjust the number of sampling points according to the learned height and width, solving the problem that the number of sampling points in conventional convolution is fixed and it is difficult to adapt to multi-scale targets. The adaptive rectangular convolution module only needs to learn two parameters, namely height and width, and has better convergence in small dataset tasks, so as to improve the adaptability of the ARDMADet model to targets of different scales and shapes. Finally, the constructed model ARDMADet significantly improves the detection accuracy of small targets while significantly reducing the number of network parameters and computational complexity.
[0015] By adopting the foregoing design scheme, the beneficial effects of the present invention are as follows: The present invention significantly improves the model's perception ability of the uncertainty of target directions and the diversity of scales by introducing the SADMA module; the SADMA module uses a mechanism that combines spatial rotation and multi-scale deformable convolution to achieve robust modeling of targets at multiple angles and multiple sizes, effectively enhancing the small target detection performance of the model in environments with uneven illumination and complex structures such as tunnels; experimental results show that in cases of occlusion, distortion, or partial contour loss, the SADMA module can still effectively extract discriminative features to avoid false detection and missed detection. In addition, the adaptive rectangular convolution module can automatically adjust the height-width ratio of the convolution kernel according to the structural characteristics of humanoid targets in tunnel scenes, which are often "slender, irregular, and with unfixed orientations", to generate an aspect-ratio adaptive rectangular receptive field, thereby more accurately extracting the structural features of fine-grained regions in the BEV map; especially in the bird's-eye view generated by lidar, due to the easy deformation or stretching of targets during the projection process, the adaptive rectangular convolution module can effectively enhance the model's adaptability to humanoid targets of different scales and different postures while controlling the computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic structural diagram of the ARDMADet model of the present invention; Figure 2 It is a schematic structural diagram of the SADMA module of the present invention; Figure 3 It is a schematic structural diagram of the deformable convolution module of the present invention; Figure 4 It is a schematic structural diagram of the Mablock module of the present invention; Figure 5 It is a schematic structural diagram of the SADMA_AR module of the present invention. Detailed implementation manners
[0017] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only partial embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0018] The terms "first", "second", "third", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0019] Tunnel anomaly detection model establishment method, establish the ARDMADet (Adaptive Rotated Deformable Multi-scale Attention Detector) model as shown in Figure 1 The ARDMADet model is an improved YOLOv10 model. The ARDMADet model includes a backbone network for feature extraction of the input image, a neck network for feature extraction and fusion of the feature map, and a head structure for detecting and classifying the fused feature map output by the neck network.
[0020] Replace the C2f module in the backbone network of the YOLOv10 model with the feature extraction unit SADMA_AR module as shown in Figure 5 The feature extraction unit SADMA_AR module is improved based on the C2f module; The feature extraction unit SADMA_AR module adds a first SADMA (Small-Aware Deformable Multi-scale Attention) module in the Concat module and the convolutional module of the C2f module, replaces both convolutional modules in the C2f module with a first Adaptive Rotated Convolution (ARConv) module. The feature map output by the first Adaptive Rotated Convolution module is sequentially subjected to normalization processing and the first activation function processing, and replaces the Bottleneck module in the C2f module with as shown in Figure 4The Mablock module shown. In this application, the adaptive rectangular convolution module is embedded into the C2f structure to replace the standard convolution module therein, so as to improve the adaptability of the model to targets of different scales and shapes. The SADMA module is connected in parallel to the C2f structure, which can further enhance the model's perception ability of the spatial region. The constructed feature extraction unit SADMA_AR module has stronger expression ability.
[0021] The Mablock module includes two second adaptive rectangular convolution modules, a second SADMA module, and a connection module connected in sequence. The feature maps output by the second adaptive rectangular convolution module are sequentially subjected to normalization processing and the first activation function processing. The connection module is used to connect the feature map input to the Mablock module with the feature map output by the second SADMA module.
[0022] The core idea of the adaptive rectangular convolution module in this embodiment is to simultaneously adjust the size of the convolution kernel in the height and width directions in a learnable manner, so that the receptive field has self-adaptability to adapt to the structural features of targets of different scales and different aspect ratios.
[0023] As Figure 2 shown, both the first SADMA module and the second SADMA module include four branches: The first branch rotates the initial feature map for the first time to obtain the first rotated feature map. In this embodiment, the first branch rotates the input C×H×W feature map counterclockwise by 90° along the H axis to obtain the W×H×C first rotated feature map; The second branch performs max-pooling and average-pooling operations on the first rotated feature map in the W dimension, and uses deformable convolution modules with different dimensions to extract features from the pooled feature maps respectively. After splicing the feature maps with different dimensions, the second activation function is used for output. In this embodiment, as Figure 3 shown, the deformable convolution modules with different dimensions refer to three deformable convolution modules with convolution kernels of 1×1, 3×3, and 7×7. The three deformable convolution modules with different dimensions form the DMC_Block module.
[0024] Among them, Z-pool focuses on max-pooling and average-pooling of the input. It can be summarized by the following formula: ; represents the input tensor C × H × W, represents dimension 0, represents the max-pooling operation, represents the average-pooling operation.
[0025] The shape tensor obtained from the input after the Z-pool is 2 × H × W, where 2 represents the dimension. Then, we introduce three parallel deformable convolutions with kernel sizes of 1, 3, and 7 respectively. This provides intermediate outputs of different scale dimensions, which are then used by the second activation function to generate attention weights. Its operation process and final output can be summarized by the following formula:
[0026]
[0027] represents the processing process of the DMC_Block module, represents the dimension, 、 and represent two-dimensional convolution operations with kernel sizes of 1, 3, and 7 respectively and a stride of 1, represents , represents the Sigmoid activation function, and represent the rotated tensor, represents the output after being processed by the activation function.
[0028] Perform an element-wise multiplication operation on the first rotated feature map and the feature map output by the second branch, and then perform the first reverse rotation and output. In this embodiment, the first reverse rotation is a 90° clockwise rotation along the H axis; The third branch performs a second rotation on the initial feature map to obtain a second rotated feature map. In this embodiment, the third branch rotates the input C×H×W feature map 90° counterclockwise along the W axis to obtain the H × C × W second rotated feature map. Perform max-pooling and average-pooling operations on the second rotated feature map in the H dimension through Z-pool. Use deformable convolution modules of different dimensions to extract features from the pooled feature maps respectively. Concatenate the feature maps of different dimensions and then output using the second activation function. Perform a second reverse rotation on the feature map output by the second activation function and then output. In this embodiment, the second reverse rotation is a 90° clockwise rotation along the W axis.
[0029] The fourth branch performs max-pooling and average-pooling operations on the initial feature map. Use deformable convolution modules of different dimensions to extract features from the pooled feature maps respectively. Concatenate the feature maps of different dimensions and then output using the second activation function. Perform an element-wise multiplication operation on the feature map output by the second activation function and the initial feature map and then output; Add the feature map output after the first reverse rotation, the feature map after the second reverse rotation, and the feature map output by the fourth branch, and then calculate the average value and output.
[0030] The SADMA module constructs multiple attention directions by rotating the input feature map in the spatial dimension to simulate the possible appearance of the target in various directions, thereby establishing a stronger direction perception ability. In each direction branch of the SADMA module, deformable convolution modules with different dimensions are used to extract local and global attention features in parallel to enhance the response ability to targets of different sizes, especially the recognition stability of small targets in complex backgrounds. Moreover, the use of deformable convolution modules enables each attention branch to have an adaptive sampling ability to adapt to features such as possible deformation and non-rigid edges of targets in the tunnel, effectively improving the modeling effect of the model on irregular and occluded targets.
[0031]
[0030] In this embodiment, the first activation function is the SILU activation function, and the second activation function is the Sigmoid activation function. In this application, the SILU activation function has a smooth and non-zero negative half-axis output, which can provide more continuous gradient information while maintaining non-linear characteristics, helping to alleviate the gradient vanishing problem and improving the stability of model training. In addition, the SILU function has a self-gating property, which can dynamically adjust the output amplitude according to the input, enhancing the network's expression ability for features of different scales. For the detection tasks of small targets and weak contrast targets involved in the present invention, SILU helps to retain more effective information, improving the detection accuracy and model robustness. Therefore, replacing the activation function from ReLU to SILU in this application not only optimizes the network training process but also improves the model's detection performance for tunnel abnormal targets.
[0032] This embodiment also provides a method for tunnel abnormal detection and early warning using the ARDMADet model established by the above tunnel abnormal detection model establishment method.
[0033]
[0031] The tunnel abnormal detection and early warning method includes the following steps executed in sequence: S1: Real-time obtain the three-dimensional point cloud data in the tunnel. This three-dimensional point cloud data is projected and converted to obtain a two-dimensional data set, that is, the three-dimensional point cloud is mapped to a two-dimensional BEV (Bird’s Eye View) plane to form an equivalent grayscale image or depth map representation, and the two-dimensional data set is preprocessed, and the preprocessed two-dimensional data set is divided into a training set, a validation set, and a test set according to a preset ratio;
[0032] In this embodiment, a conventional lidar device is used to collect three-dimensional point cloud data. The lidar device can be installed in the tunnel or on the rail vehicle. The preprocessing in step S1 includes sequentially performing background filtering, ROI region cropping, and noise suppression on the two-dimensional data set.
[0034] S2: Input the training set into the ARDMADet model established by using the tunnel anomaly detection model establishment method described in any of the above items to train the ARDMADet model; In this application, Jetson Nano is used as the edge computing platform. After the trained ARDMADet model is optimized in structure and quantized and compressed, it is converted into the form of a TensorRT engine and deployed on the Jetson Nano platform. Jetson Nano has strong edge inference capabilities and can run the object detection model at a high frame rate under low power consumption conditions to meet the real-time detection requirements.
[0035] S3: Use the WGIoU loss function to perform backward adjustment on the trained ARDMADet model to obtain the adjusted ARDMADet model. The calculation formula of the WGIoU loss function is as follows: ; where A represents the ground truth box, B represents the preset bounding box, C represents the smallest box containing the ground truth box and the predicted box, represents the coordinates of the center of the anchor box, represents the coordinates of the center of the target box, and represent the size of the minimum enclosing box. The entire loss function consists of three parts: the Generalized IoU term , which is used to measure the overall overlap degree between the predicted box and the ground truth box, and remains effective gradients especially when IoU is 0. The center point penalty term , which penalizes the offset between the center of the predicted box and the center of the ground truth box in the form of Gaussian decay to improve the accuracy and stability of bounding box localization.
[0036] The non-monotonic focusing coefficient r is defined as: where , is defined as the outlier degree, which is used to describe the quality of the anchor box, is the IoU threshold between the predicted bounding box and the ground truth bounding box, is The moving average of can be dynamically adjusted during training to adjust the overall gain. A small outlier degree means high-quality anchor boxes, and a small gradient gain is assigned to them, so that the bounding box regression focuses on the anchor boxes of ordinary quality. Assigning a smaller gradient gain to the anchor boxes with a larger outlier degree will effectively prevent low-quality image data from generating large harmful gradients and affecting the detection quality. and are hyperparameters and can be appropriately adjusted according to different models. That is to use A non - monotonic focusing coefficient is constructed to increase the penalty for high - IoU samples and reduce the attention to low - IoU samples, so that the network focuses more on high - quality candidate boxes.
[0037] In a complex tunnel environment, using the WGIoU loss function to perform back - adjustment on the ARDMADet model can significantly improve the robustness and accuracy of detection. This is mainly due to its three core mechanisms: 1. The generalized IoU term effectively solves the problems of inherent object sparsity, boundary ambiguity, and occlusion in BEV projection (especially at the far end or corners of the tunnel). It can provide effective gradients even when the predicted box and the ground - truth box have zero overlap, significantly reducing missed detections; 2. The center - point penalty term based on Gaussian decay strengthens the precise fine - tuning of target center localization and noise suppression, improving the accuracy and inter - frame stability of position estimation for targets such as personnel under the influence of point - cloud noise; 3. The key non - monotonic focusing coefficient r (based on the outlier degree β and hyperparameters and ) dynamically adjusts the training weights, intelligently suppresses the interference of noise / outlier samples such as "ghost targets" generated by tunnel - wall reflections, false alarms of fixed facilities, and low - quality sparse targets, while focusing on high - quality learnable samples, thus significantly reducing the environmental false - detection rate. These mechanisms work together to enable the model to obtain more reliable target - detection performance in complex tunnel scenarios (such as occlusion, structural interference, and point - cloud sparsity), providing a more accurate environmental perception basis for the safety - warning system.
[0038] S4: Input the test set into the adjusted ARDMADet model for testing to obtain the test results; the test results are encapsulated in JSON format, including fields such as target category, confidence, center coordinates, and bounding - box size, and are sent to the edge gateway in the local network through the lightweight MQTT protocol. The entire system can achieve a complete closed - loop inference process without relying on an external server, with advantages such as low latency, low power consumption, and high adaptability, and is suitable for deployment in resource - constrained industrial sites.
[0039] S5: If the test results show that there are people in the preset area, an alarm prompt is issued; the alarm prompt here includes but is not limited to emitting a warning sound, sending a warning message to the management personnel, voice alarm, etc.
[0040] In this application, a computer product that can implement the above - mentioned tunnel anomaly detection and warning method is also provided.
[0041] A computer product includes a computer program, which when executed by a processor implements the tunnel anomaly detection and warning method described in any one of the above.
[0042] In this application, a storage medium storing a program that can implement the above - mentioned tunnel anomaly detection and warning method is also provided.
[0043] A computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, the tunnel anomaly detection and early warning method described in any one of the above is implemented.
[0044] In summary, by introducing the SADMA module, the present invention significantly improves the model's perception ability of the uncertainty of the target direction and the diversity of scales; the SADMA module uses a mechanism that combines spatial rotation and multi-scale deformable convolution to achieve robust modeling of the target at multiple angles and multiple sizes, effectively enhancing the small target detection performance of the model in environments with uneven lighting and complex structures such as tunnels; experimental results show that in the case of occlusion, distortion, or partial contour loss, the SADMA module can still effectively extract discriminative features to avoid false detection and missed detection; In addition, the adaptive rectangular convolution module can automatically adjust the height and width ratio of the convolution kernel according to the structural characteristics of humanoid targets in tunnel scenes, which are often "slender, irregular, and with unfixed orientations", to generate an aspect ratio adaptive rectangular receptive field, thereby more accurately extracting the structural features of fine-grained regions in the BEV map; especially in the bird's-eye view map generated by lidar, due to the easy deformation or stretching of the target during the projection process, the adaptive rectangular convolution module can effectively enhance the model's adaptability to humanoid targets of different scales and different poses, while controlling the computational complexity and maintaining efficient operation on edge platforms such as Jetson Nano. At the same time, to address the common problems in tunnel BEV maps such as low gray values, weak target contrast, and strong background interference, the SILU activation function is used to replace the traditional ReLU. The SILU function still retains a non-zero response in the negative input interval, has better information retention ability and gradient smoothness, which helps to enhance the discriminative power of the network for dark and weak target regions and alleviate the gradient vanishing problem. This mechanism can significantly improve the detection accuracy and robustness of small humanoid targets in low-contrast scenarios in the BEV input used in this application.
[0045] Finally, in a complex tunnel environment, the WGIoU loss function is used to perform backward adjustment on the ARDMADet model, which can significantly improve the robustness and accuracy of detection.
[0046] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. Method for establishing tunnel anomaly detection model, characterized by: Build the ARDMADet model, which is an improved YOLOv10 model. The ARDMADet model includes a backbone network for feature extraction of the input image, a neck network for feature extraction and fusion of the feature map, and a head structure for detecting and classifying the fused feature map output by the neck network; Replace the C2f module in the backbone network of the YOLOv10 model with the feature extraction unit SADMA_AR module, which is improved based on the C2f module; The feature extraction unit SADMA_AR module adds a first SADMA module in the Concat module and the convolutional module of the C2f module, replaces the two convolutional modules in the C2f module with the first adaptive rectangular convolutional module. The feature map output by the first adaptive rectangular convolutional module is sequentially processed by normalization and the first activation function, and replaces the Bottleneck module in the C2f module with the Mablock module; The Mablock module includes two second adaptive rectangular convolutional modules, a second SADMA module, and a connection module connected in sequence. The feature map output by the second adaptive rectangular convolutional module is sequentially processed by normalization and the first activation function. The connection module is used to connect the feature map input to the Mablock module with the feature map output by the second SADMA module; 2. The method for establishing a tunnel anomaly detection model according to claim 1, characterized in that: Both the first SADMA module and the second SADMA module include four branches: The first branch rotates the initial feature map for the first time to obtain the first rotated feature map; The second branch performs max-pooling and average-pooling operations on the first rotated feature map, uses deformable convolutional modules with different dimensions to extract features from the pooled feature maps respectively, splices the feature maps with different dimensions, and outputs after passing through the second activation function; Perform an element-wise multiplication operation on the first rotated feature map and the feature map output by the second branch, and then perform the first reverse rotation and output; The third branch rotates the initial feature map for the second time to obtain the second rotated feature map, performs max-pooling and average-pooling operations on the second rotated feature map, uses deformable convolutional modules with different dimensions to extract features from the pooled feature maps respectively, splices the feature maps with different dimensions, and outputs after passing through the second activation function, and outputs after performing the second reverse rotation on the feature map output by the second activation function; The fourth branch performs max-pooling and average-pooling operations on the initial feature map, uses deformable convolutional modules with different dimensions to extract features from the pooled feature maps respectively, splices the feature maps with different dimensions, and outputs after passing through the second activation function, and outputs after performing an element-wise multiplication operation on the feature map output by the second activation function and the initial feature map; Add and average the feature map output after the first reverse rotation, the feature map output after the second reverse rotation, and the feature map output by the fourth branch and then output.
3. The method for establishing a tunnel anomaly detection model according to claim 2, characterized in that: The first rotation is a 90° counterclockwise rotation along the H axis, the first reverse rotation is a 90° clockwise rotation along the H axis, the second rotation is a 90° counterclockwise rotation along the W axis, and the second reverse rotation is a 90° clockwise rotation along the W axis.
4. The method for establishing a tunnel anomaly detection model according to claim 3, wherein: The second activation function is the Sigmoid activation function.
5. The method for establishing a tunnel anomaly detection model according to claim 1, wherein: The first activation function is the SILU activation function.
6. Tunnel anomaly detection and early warning method, characterized in that: It includes the following steps executed in sequence: S1: Real-time obtain the three-dimensional point cloud data in the tunnel. The three-dimensional point cloud data is projected and converted to obtain a two-dimensional data set, and the two-dimensional data set is preprocessed, and the preprocessed two-dimensional data set is divided into a training set, a validation set, and a test set according to a preset ratio; S2: Input the training set into the ARDMADet model established by using the tunnel anomaly detection model establishment method described in any one of the above claims 1-5 to train the ARDMADet model; S3: Use the WGIoU loss function to perform backward adjustment on the trained ARDMADet model to obtain the adjusted ARDMADet model. The calculation formula of the WGIoU loss function is as follows: ; Among them, A represents the ground truth box, B represents the preset bounding box, and C represents the smallest box containing the ground truth box and the predicted box. represents the coordinates of the center of the anchor box, represents the coordinates of the center of the target box, and represents the size of the minimum bounding box, , where , is the outlier degree, is the IoU threshold between the predicted bounding box and the ground truth bounding box, is 's moving average. A small outlier degree means a high-quality anchor box. and are hyperparameters, is a non-monotonic focusing coefficient constructed using . S4: Input the test set into the adjusted ARDMADet model for testing to obtain the test result; S5: If the test result shows that there are people in the preset area, an alarm prompt is issued.
7. The tunnel anomaly detection and early warning method according to claim 6, characterized in that: The preprocessing in step S1 includes sequentially performing background filtering, ROI region cropping, and noise suppression on the two-dimensional data set.
8. A computer product, characterized in that: It includes a computer program that, when executed by a processor, implements the tunnel anomaly detection and warning method described in any one of the above claims 6-7.
9. A computer storage medium, on which a computer program is stored, characterized in that: When the computer program is executed by a processor, it implements the tunnel anomaly detection and warning method described in any one of the above claims 6-7.
Citation Information
Patent Citations
Salient target detection method and system, storage medium and product
CN118968041A
Improved YOLO v8-based mushroom stick cultivation mushroom grading method and system
CN119206711A
Tunnel disease detection method based on anchor-frame-free adaptive convolutional network and inspection robot
CN119810053A
Face target detection method based on improved YOLOv8
CN120088835A
PCB surface defect detection system and method based on improved YOLOv10
CN120125585A