Dam global hidden danger detection method based on radar image interpretation and visual large model
By integrating 3D ground-penetrating radar imaging with visual annotation and employing RFE-Attention and CS-VIT-based detection methods, the problems of low efficiency and insufficient feature extraction adaptability in dam hazard detection have been solved, achieving high-precision and stable hazard location and risk classification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI TRANSPORT CONSULTING & DESIGN INST
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies for detecting hidden dangers in dams suffer from low efficiency, high subjectivity, and destructiveness to the dam body. Furthermore, general detection networks lack adaptability for feature extraction in 3D radar imaging scenarios, making it difficult to achieve full-category coverage, especially in detecting small sample categories.
By integrating 3D ground-penetrating radar imaging and visual annotation, using RFE-Attention to enhance hazard features, and combining CS-VIT-specific detection and cross-domain pseudo-label optimization, we can achieve the location and risk classification of dam hazards with fewer annotations and across regions.
It improved detection accuracy and generalization robustness, reduced the false negative and false positive rates, enhanced the detection coverage of niche defects and the overall detection stability, and generated a reliable three-dimensional hazard distribution map of the dam body.
Smart Images

Figure CN122023341A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent inspection and hidden danger detection technology for dikes, and in particular to a method for detecting hidden dangers across the entire area of dikes based on radar image interpretation and visual large model. Background Technology
[0002] As the core infrastructure of flood control and water supply projects, the early identification of surface and deep hidden dangers in dikes is of great significance to the safety of the project. Existing inspection methods mainly include manual inspection combined with ground penetrating radar and drilling verification, as well as automated detection systems for single modes. Manual inspection relies on personnel experience to observe surface anomalies such as cracks on the dam surface, and uses handheld ground penetrating radar to conduct deep detection, and then verifies through drilling. Although it can obtain a certain degree of reliability, it has problems such as low efficiency, strong subjectivity, and damage to the dam body, and it is difficult to meet the high-frequency inspection needs of long-distance dike sections.
[0003] With the development of target detection models and large-scale vision models, existing technologies have begun to attempt to use general detection networks for radar images and engineering structure defect identification. However, significant shortcomings still exist in the 3D radar imaging scenario of dams. 3D radar imaging data contains physical features such as reflection intensity, phase difference, and layered structure. Its data distribution differs significantly from that of natural images. General convolutional feature extraction and standard attention mechanisms are difficult to highlight the differentiated features between defects and the background, resulting in insufficient feature extraction adaptability and missed detections. On the other hand, dam defect annotation usually relies on professionals to complete the work by combining radar interpretation and on-site verification. This is costly and time-consuming, resulting in scarce annotated samples and an unbalanced class distribution. General multi-class shared training frameworks are prone to class bias and have insufficient detection capabilities for small sample classes, making it difficult to achieve full class coverage.
[0004] Therefore, how to provide a method for detecting hidden dangers in the entire dam area based on radar image interpretation and visual large model is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose a method for detecting hidden dangers in dams across the entire region based on radar image interpretation and a large visual model. This invention integrates three-dimensional ground-penetrating radar imaging and visual annotation, uses RFE-Attention to enhance the features of the dam, and combines CS-VIT-specific detection and cross-domain pseudo-label optimization to achieve dam hidden danger location and risk classification with fewer annotations and across regions, and effectively improves detection accuracy and generalization robustness.
[0006] The method for detecting hidden dangers in the entire dam area based on radar image interpretation and visual large model according to an embodiment of the present invention includes the following steps:
[0007] Data on the embankment was collected using a mobile inspection platform equipped with a 3D ground-penetrating radar, resulting in 3D radar imaging data, visual images, and the corresponding GPS coordinates of the embankment.
[0008] The 3D radar imaging data is denoised and then registered with the GPS coordinates of the dam to obtain a standardized radar feature map.
[0009] The location information of surface diseases is obtained based on visual images and spatially aligned with standardized radar feature maps to form standardized multi-mode feature maps.
[0010] The standardized multi-mode feature map is input into the radar feature enhancement attention module RFE-Attention, and feature enhancement processing is performed through the feature decomposition unit, the dual-branch weight generator and the weight fusion unit to output the enhanced feature map.
[0011] The enhanced feature map input class-specific detection architecture CS-VIT is used to extract features through a shared feature extraction layer, class-specific training branches and cross-branch fusion layers to obtain multi-class disease detection results;
[0012] Cross-domain adaptive training and optimization were performed on CS-VIT, and the detection results of multiple diseases were updated based on the optimized CS-VIT to obtain the updated detection results of multiple diseases.
[0013] The updated multi-category disease detection results are correlated with the dam's GPS coordinates and surface disease location marking information to generate a three-dimensional hazard distribution map of the dam body.
[0014] Optionally, obtaining the three-dimensional radar imaging data, visual images, and corresponding GPS coordinates of the dam specifically includes:
[0015] Based on the 3D radar data acquisition module, a mobile inspection platform equipped with 3D ground-penetrating radar is provided. It moves along the dam inspection path and collects data during the movement. During the acquisition process, visual images and the GPS coordinates of the dam corresponding to the acquisition location are acquired simultaneously.
[0016] The three-dimensional ground-penetrating radar operates at a center frequency of 200MHz to 500MHz to conduct three-dimensional detection of the dam body within a depth range of 0m to 5m, obtaining three-dimensional radar imaging data. The three-dimensional radar imaging data is output as a point cloud to show the spatial scattering distribution of the dam body, and as a grayscale profile map to show the radar profile imaging results along the inspection path.
[0017] Optionally, obtaining the standardized radar feature map specifically includes:
[0018] Wavelet thresholding is performed on the 3D radar imaging data to eliminate electromagnetic interference. Wavelet decomposition is performed and the coefficients obtained by decomposition are thresholded. Inverse wavelet reconstruction is performed on the thresholded coefficients, and the denoised 3D radar imaging data is output.
[0019] The denoised 3D radar imaging data is subjected to coordinate registration processing to determine the correspondence between each acquisition location in the denoised 3D radar imaging data and the GPS coordinates of the dam, and the data is bound together and the bound radar data is output.
[0020] The bound radar data is standardized by unifying it with a preset coordinate reference and a preset data format, and the standardized result is output as a standardized radar feature map.
[0021] Optionally, the formation of the standardized multimodal feature map specifically includes:
[0022] The surface disease is marked on the visual image. The marking includes the marking position, marking length and marking width of the surface disease, and the surface disease location marking information is output.
[0023] Based on the GPS coordinates of the dam, establish the correspondence between visual images and standardized radar feature maps at the same acquisition location;
[0024] Based on the correspondence, the surface disease location annotation information is mapped to the spatial range of the standardized radar feature map to obtain the mapped surface disease location annotation information, which is then fused with the standardized radar feature map to output a standardized multi-mode feature map.
[0025] Optionally, the output of the enhanced feature map specifically includes
[0026] The standardized multi-mode feature map is input into the radar feature enhancement attention module RFE-Attention, which includes a feature decomposition unit, a dual-branch weight generator and a weight fusion unit, and extracts radar channel, i.e., reflection intensity data, from the standardized multi-mode feature map as data to be processed.
[0027] The feature decomposition unit performs feature decomposition on the reflection intensity data. A three-to-three window is used to perform sliding calculation on the reflection intensity data. The variance of the reflection value within the window is used as the local scattering feature, and the phase difference of the reflected wave along the depth direction is used as the global layering feature.
[0028] The dual-branch weight generator performs weight mapping processing on the local scattering features and the global hierarchical features respectively. The local scattering features are mapped by the Sigmoid function to output local attention weights, and the global hierarchical features are mapped by the Softmax function to output hierarchical attention weights. The local attention weights and hierarchical attention weights are multiplied to obtain the RFE-Attention weight map.
[0029] The weight fusion unit multiplies the RFE-Attention weight map pixel by pixel with the standardized multimodal feature map and outputs the enhanced feature map.
[0030] Optionally, obtaining the multi-category disease detection results specifically includes:
[0031] The enhanced feature map is input into the class-specific detection architecture CS-VIT, which is a large visual model including a shared feature extraction layer, class-specific training branches set according to the number of disease categories, and a cross-branch fusion layer. The class-specific training branches include a VIT encoder, an existence judgment head, and a target box regression head.
[0032] In the shared feature extraction layer, convolution and pooling are performed sequentially on the enhanced feature map. Through four layers of convolution and pooling operations, the feature channels are compressed into 256 channels to obtain the shared feature map.
[0033] In each dedicated training branch, the shared feature map is divided into blocks according to the preset block division rules, forming a block sequence, so that the block sequence contains sixty-four blocks.
[0034] In the VIT encoder, the tile sequence is encoded using a six-layer Transformer to obtain the class-specific features corresponding to the current disease category;
[0035] In the existence judgment head, the class-specific features are processed sequentially with fully connected processing and Sigmoid activation processing to obtain the judgment result of whether the current disease category exists, and the detection confidence corresponding to the judgment result is obtained.
[0036] In the target bounding box regression head, the class-specific features are processed four times in succession to obtain the regression results used to determine the upper left and lower right corner positions of the target bounding box, forming the branch detection results of the current disease category;
[0037] In the cross-branch fusion layer, the branch detection results output by various dedicated training branches are subjected to weighted voting fusion processing. The results are weighted according to the detection confidence and the optimal result is selected. The detection boxes and category labels of all categories of diseases are output as multi-category disease detection results.
[0038] Optionally, obtaining the updated multi-category disease detection results specifically includes:
[0039] Read the detection results of multiple types of diseases, and identify the three-dimensional radar imaging data with disease category and target box annotation information as labeled radar data, otherwise identify it as unlabeled radar data. Then divide the unlabeled radar data into different regional data sets according to the regional source to form the input data set for cross-domain adaptive training and optimization.
[0040] Perform cross-domain data augmentation processing, call the domain adaptation network to migrate and align the radar imaging distribution of different regional datasets, and output the aligned unlabeled radar data as cross-domain augmentation data;
[0041] The radar physical enhancement sample generation process is executed. The radar physical enhancement operator is called to enhance the labeled radar data and cross-domain enhancement data. The reflection intensity attenuation is adjusted in the propagation difference simulation process, and regional specific noise is superimposed to output regional adaptive enhancement samples.
[0042] Perform dynamic weight loss training preparation processing, set training loss weights corresponding to the current disease class for each disease class, and introduce constant terms to avoid training loss weights being zero.
[0043] Dynamic weight loss training is performed. During the training process, the sample frequency of each disease category is counted in real time, and the detection confidence of each disease category is counted in real time. The loss weights corresponding to each disease category are updated. The updated loss weights are used to weight the classification loss and regression loss of CS-VIT. At the same time, the model training is updated. The labeled radar data and the regional adaptive enhancement samples are used as training data input to obtain the CS-VIT model with phased optimization.
[0044] Semi-supervised pseudo-label iterative training is performed. The CS-VIT model with phased optimization is used to detect cross-domain augmented data. The detection results are filtered according to the detection confidence. The detection results with a detection confidence of not less than 0.85 are selected as pseudo-labels. The labeled radar data and pseudo-label data are mixed to form an iterative training set. The CS-VIT model is iteratively trained and updated. The pseudo-labels are updated after each iteration of training to obtain the optimized CS-VIT model.
[0045] The optimized CS-VIT model is used to update the detection results of multiple disease categories, and cross-branch fusion processing is performed to output the updated detection results of multiple disease categories.
[0046] Optionally, the generation of the three-dimensional hazard distribution map of the dam body specifically includes:
[0047] The updated multi-category disease detection results are processed by coordinate association. The spatial location corresponding to each disease detection result is bound to the GPS coordinates of the dam and written into the data record corresponding to the GPS coordinates of the dam, forming a set of disease detection records with coordinates.
[0048] Simultaneously perform coordinate association processing on the surface disease location annotation information to form a set of surface disease annotation records with coordinates;
[0049] Based on the coordinate-based disease detection records and coordinate-based surface disease location annotation information, a three-dimensional aggregation process is performed. The disease category, target box and detection confidence corresponding to the same dam GPS coordinates are aggregated. The surface disease annotation location, annotation length and annotation width corresponding to the same dam GPS coordinates are aggregated. The aggregated surface disease information and deep disease information are distributed and organized in three dimensions according to the spatial location of the dam to obtain a three-dimensional hidden danger information set.
[0050] The three-dimensional hazard information set is processed using CT-style output. The annotation position, annotation length and annotation width of the surface disease are marked by visual image and written into the corresponding spatial position. The annotation depth, annotation volume and radar reflection characteristics of the deep disease are written into the corresponding spatial position. The surface cracks are judged to be connected to the deep seepage channels by multi-mode data association and written into the corresponding spatial position. The location, depth and risk level information are written into the corresponding spatial position to generate a three-dimensional hazard distribution map of the dam body.
[0051] The beneficial effects of this invention are:
[0052] This invention incorporates radar reflection intensity and phase difference information into attention weight calculation through RFE-Attention, which enhances and suppresses local scattering features and global layered features, reduces the confusion between defects and background and the interference of interlayer noise, and improves the adaptability of 3D radar imaging feature extraction and the ability to characterize hidden dangers.
[0053] This invention adopts a CS-VIT class-specific detection architecture, which uses a shared feature extraction layer in conjunction with class-specific training branches set according to disease categories to realize existence judgment and target box regression. The detection results of all categories are output through weighted voting by a cross-branch fusion layer, thereby alleviating the class bias caused by small sample size and class imbalance, and improving the detection coverage of niche diseases and the overall detection stability.
[0054] This invention forms a detection-update closed loop through cross-domain adaptive training and optimization. It combines domain adaptation transfer, radar physical enhancement, dynamic weight loss and pseudo-label iterative training to improve the generalization robustness of the model in environments with few labels and multiple regions. Finally, it outputs the updated detection results and associates them with GPS and surface labeling information to generate a three-dimensional hazard distribution map of the dam body, providing location, depth and risk level information for repair. Attached Figure Description
[0055] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0056] Figure 1 This is a flowchart of the method for detecting hidden dangers in the entire dam area based on radar image interpretation and visual large model proposed in this invention;
[0057] Figure 2 This is a schematic diagram of the CS-VIT-based dedicated detection architecture for the dam full-area hidden danger detection method based on radar image interpretation and visual large model proposed in this invention.
[0058] Figure 3 This is a schematic diagram of the RFE-Attention module structure of the dam full-area hidden danger detection method based on radar image interpretation and visual large model proposed in this invention. Detailed Implementation
[0059] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0060] refer to Figures 1-3 A method for detecting hidden dangers in the entire dam area based on radar image interpretation and visual large model includes the following steps:
[0061] A mobile inspection platform equipped with a three-dimensional ground-penetrating radar is used to collect dam data, obtain three-dimensional radar imaging data, visual images and corresponding GPS coordinates of the dam. The three-dimensional radar imaging data includes point cloud data and grayscale profile data.
[0062] The 3D radar imaging data is denoised and then registered with the GPS coordinates of the dam to obtain a standardized radar feature map.
[0063] The location annotation information of surface diseases is obtained based on visual images and spatially aligned with standardized radar feature maps to form standardized multi-mode feature maps.
[0064] The standardized multi-mode feature map is input into the radar feature enhancement attention module RFE-Attention, and feature enhancement processing is performed through the feature decomposition unit, the dual-branch weight generator and the weight fusion unit to output the enhanced feature map.
[0065] The enhanced feature map input class-specific detection architecture CS-VIT is used to extract features through a shared feature extraction layer, class-specific training branches and cross-branch fusion layers to obtain multi-class disease detection results;
[0066] Cross-domain adaptive training and optimization are performed on CS-VIT. The cross-domain adaptive training and optimization sequentially performs cross-domain data augmentation, radar physical augmentation sample generation, dynamic weight loss training, and semi-supervised pseudo-label iterative training. Based on the optimized CS-VIT, the detection results of multiple diseases are updated to obtain the updated multi-category disease detection results.
[0067] The updated multi-category disease detection results are correlated with the dam's GPS coordinates and surface disease location marking information to generate a three-dimensional hazard distribution map of the dam body.
[0068] In this embodiment, obtaining the three-dimensional radar imaging data, visual images, and corresponding GPS coordinates of the dam specifically includes:
[0069] Based on the 3D radar data acquisition module, a mobile inspection platform equipped with 3D ground-penetrating radar is provided. It moves along the dam inspection path and collects data during the movement. During the acquisition process, visual images and the GPS coordinates of the dam corresponding to the acquisition location are acquired simultaneously.
[0070] The three-dimensional ground-penetrating radar operates at a center frequency of 200MHz to 500MHz to conduct three-dimensional detection of the dam body within a depth range of 0m to 5m, obtaining three-dimensional radar imaging data. The three-dimensional radar imaging data is output as a point cloud to show the spatial scattering distribution of the dam body, and as a grayscale profile map to show the radar profile imaging results along the inspection path.
[0071] In this embodiment, obtaining the standardized radar feature map specifically includes:
[0072] Wavelet thresholding is performed on the 3D radar imaging data to eliminate electromagnetic interference. Wavelet decomposition is performed and the coefficients obtained by decomposition are thresholded. Inverse wavelet reconstruction is performed on the thresholded coefficients, and the denoised 3D radar imaging data is output.
[0073] The denoised 3D radar imaging data is subjected to coordinate registration processing to determine the correspondence between each acquisition location in the denoised 3D radar imaging data and the GPS coordinates of the dam, and the data is bound together and the bound radar data is output.
[0074] The bound radar data is standardized by unifying it with a preset coordinate reference and a preset data format, and the standardized result is output as a standardized radar feature map.
[0075] In this embodiment, the formation of the standardized multimodal feature map specifically includes:
[0076] The surface disease is marked on the visual image. The marking includes the marking position, marking length and marking width of the surface disease, and the surface disease location marking information is output.
[0077] Based on the GPS coordinates of the dam, establish the correspondence between visual images and standardized radar feature maps at the same acquisition location;
[0078] Based on the correspondence, the surface disease location annotation information is mapped to the spatial range of the standardized radar feature map to obtain the mapped surface disease location annotation information, which is then fused with the standardized radar feature map to output a standardized multi-mode feature map. The standardized multi-mode feature map includes radar channels, i.e., reflection intensity data.
[0079] In this embodiment, the output of the enhanced feature map specifically includes
[0080] The standardized multi-mode feature map is input into the radar feature enhancement attention module RFE-Attention, which includes a feature decomposition unit, a dual-branch weight generator and a weight fusion unit, and extracts radar channel, i.e., reflection intensity data, from the standardized multi-mode feature map as data to be processed.
[0081] The feature decomposition unit performs feature decomposition on the reflection intensity data. A three-to-three window is used to perform sliding calculation on the reflection intensity data, and the variance of the reflection value within the window is used as the local scattering feature. The local scattering feature can reflect the local reflection anomaly of the disease. At the same time, the phase difference of the reflected wave along the depth direction is used as the global layering feature. The global layering feature can avoid local errors caused by the structural characteristics of the dam.
[0082] A dual-branch weight generator performs weight mapping processing on local scattering features and global hierarchical features respectively. The local scattering features are mapped by the Sigmoid function to output local attention weights, and the global hierarchical features are mapped by the Softmax function to output hierarchical attention weights. The local attention weights and hierarchical attention weights are multiplied to obtain the RFE-Attention weight map. The larger the value of the local attention weight, the more likely the corresponding area is to be a disease. The hierarchical attention weight is used to highlight the depth layer where the disease is located and suppress noise in other layers.
[0083] The weight fusion unit multiplies the RFE-Attention weight map with the standardized multimodal feature map pixel by pixel. The weight value at each pixel position is multiplied with the feature value at the same pixel position in the standardized multimodal feature map. The resulting product value is written to the corresponding pixel position in the enhanced feature map. When the standardized multimodal feature map contains different channels, the weight value at the same pixel position is multiplied with the feature value at the corresponding pixel position of each channel and written to the corresponding channel position in the enhanced feature map, and the enhanced feature map is output.
[0084] This invention incorporates radar channel reflection intensity and phase difference information along the depth direction into attention calculation through RFE-Attention. It utilizes local scattering features and global layered features to adaptively increase the weight of the disease area and suppress inter-layer noise. This enhances the disease features in the multi-mode feature map and weakens background interference, thereby improving the distinguishability and detection stability of 3D radar imaging for hidden dangers such as piping and voids, reducing missed detections and false detections, and improving robustness under different soil imaging conditions. This provides more reliable input features for subsequent class-specific detection and 3D hidden danger distribution representation.
[0085] In this embodiment, obtaining the detection results of the multiple categories of diseases specifically includes:
[0086] The enhanced feature map is input into the class-specific detection architecture CS-VIT, which is a large visual model including a shared feature extraction layer, class-specific training branches set according to the number of disease categories, and a cross-branch fusion layer. The class-specific training branches include a VIT encoder, an existence judgment head, and a target box regression head.
[0087] In the shared feature extraction layer, convolution and pooling are performed sequentially on the enhanced feature map. Through four layers of convolution and pooling operations, the feature channels are compressed into 256 channels to obtain the shared feature map.
[0088] In each dedicated training branch, the shared feature map is divided into blocks according to the preset block division rules, forming a block sequence, so that the block sequence contains sixty-four blocks.
[0089] In the VIT encoder, the tile sequence is encoded using a six-layer Transformer to obtain the class-specific features corresponding to the current disease category;
[0090] In the existence judgment head, the class-specific features are processed sequentially with fully connected processing and Sigmoid activation processing to obtain the judgment result of whether the current disease category exists, and the detection confidence corresponding to the judgment result is obtained.
[0091] In the target bounding box regression head, the class-specific features are processed four times in succession to obtain the regression results used to determine the upper left and lower right corner positions of the target bounding box, forming the branch detection results of the current disease category;
[0092] In the cross-branch fusion layer, the branch detection results output by various dedicated training branches are subjected to weighted voting fusion processing. The results are weighted according to the detection confidence and the optimal result is selected. The detection boxes and category labels of all categories of diseases are output as multi-category disease detection results.
[0093] This invention employs a CS-VIT class-specific detection architecture, which extracts stable high-dimensional representations from enhanced feature maps through shared feature extraction, and sets up independent training branches for each disease category to learn class-specific features separately. Combining existence judgment and target box regression, it achieves accurate localization under small sample conditions. Then, the cross-branch fusion layer outputs the detection results of all categories based on confidence-weighted voting, thereby significantly alleviating the class bias caused by class imbalance, improving the detection rate of niche diseases and the overall detection consistency, enhancing robustness to complex radar imaging scenarios, and reducing false positives and false negatives.
[0094] In this embodiment, obtaining the updated multi-category disease detection results specifically includes:
[0095] Read the detection results of multiple types of diseases, and identify the three-dimensional radar imaging data with disease category and target box annotation information as labeled radar data, otherwise identify it as unlabeled radar data. Then divide the unlabeled radar data into different regional data sets according to the regional source to form the input data set for cross-domain adaptive training and optimization.
[0096] Cross-domain data augmentation processing is performed by calling the domain adaptation network to migrate and align the radar imaging distribution of different regional datasets. The aligned unlabeled radar data is then output as cross-domain augmented data to expand the amount of training data and achieve consistency between the imaging distribution of different regions and the imaging distribution of labeled data.
[0097] The radar physical enhancement sample generation process is performed by calling the radar physical enhancement operator to enhance the labeled radar data and cross-domain enhancement data. The radar physical enhancement operator sets the radar wave propagation difference parameters according to different soil conditions and performs propagation difference simulation processing on the radar imaging data. In the propagation difference simulation processing, the reflection intensity attenuation is adjusted and regional noise is superimposed to output regional adaptive enhancement samples, so as to generate regional adaptive enhancement samples and improve the robustness of the model to different environments.
[0098] The dynamic weight loss training preparation process is performed. For each disease class, a training loss weight corresponding to the current disease class is set. The training loss weight increases as the frequency of disease class samples decreases and as the detection confidence of disease class decreases. A constant term is introduced to avoid the training loss weight being zero. The constant term is set to 0.1 to achieve dynamic weighting of class frequency and detection confidence and alleviate class bias caused by class imbalance.
[0099] Dynamic weight loss training is performed. During the training process, the sample frequency of each disease category is counted in real time, and the detection confidence of each disease category is counted in real time. The loss weights corresponding to each disease category are updated. The updated loss weights are used to weight the classification loss and regression loss of CS-VIT. At the same time, the model training is updated. The labeled radar data and the regional adaptive enhancement samples are used as training data input to obtain the CS-VIT model with phased optimization.
[0100] Semi-supervised pseudo-label iterative training is performed. The CS-VIT model with phased optimization is used to detect cross-domain augmented data. The detection results are filtered according to the detection confidence. The detection results with a detection confidence of not less than 0.85 are selected as pseudo-labels. The labeled radar data and pseudo-label data are mixed to form an iterative training set. The CS-VIT model is iteratively trained and updated. The pseudo-labels are updated after each iteration of training to obtain the optimized CS-VIT model.
[0101] The optimized CS-VIT model is used to update the detection results of multiple disease categories, and cross-branch fusion processing is performed to output the updated detection results of multiple disease categories.
[0102] This invention expands the effective training samples by combining cross-domain imaging distribution migration alignment with radar physical enhancement, and suppresses class bias by using a dual-factor dynamic weighting of class frequency and detection confidence. At the same time, it uses high-confidence pseudo-labels to iteratively mine the value of unlabeled data, and achieves adaptive optimization of CS-VIT under conditions of few labels, cross-regional soil differences and noise variations. This improves the stability, generalization robustness and accuracy of multi-category disease detection results and reduces false negatives and false positives.
[0103] In this embodiment, the generation of the three-dimensional hidden danger distribution map of the dam body specifically includes:
[0104] The updated multi-category disease detection results are processed by coordinate association. The spatial location corresponding to each disease detection result is bound to the GPS coordinates of the dam and written into the data record corresponding to the GPS coordinates of the dam, forming a set of disease detection records with coordinates.
[0105] Simultaneously perform coordinate association processing on the surface disease location annotation information to form a set of surface disease annotation records with coordinates;
[0106] Based on the coordinate-based disease detection records and coordinate-based surface disease location annotation information, a three-dimensional aggregation process is performed. The disease category, target box and detection confidence corresponding to the same dam GPS coordinates are aggregated. The surface disease annotation location, annotation length and annotation width corresponding to the same dam GPS coordinates are aggregated. The aggregated surface disease information and deep disease information are distributed and organized in three dimensions according to the spatial location of the dam to obtain a three-dimensional hidden danger information set.
[0107] The three-dimensional hazard information set is processed using CT-style output. The annotation position, annotation length and annotation width of the surface disease are marked by visual image and written into the corresponding spatial position. The annotation depth, annotation volume and radar reflection characteristics of the deep disease are written into the corresponding spatial position. The surface cracks are judged to be connected to the deep seepage channels by multi-mode data association and written into the corresponding spatial position. The location, depth and risk level information are written into the corresponding spatial position to generate a three-dimensional hazard distribution map of the dam body.
[0108] This invention unifies and three-dimensionally summarizes the updated multi-category disease detection results with the dam's GPS coordinates and surface disease location marking information. It integrates disease categories, target boxes, detection confidence levels, and the length and width information of surface markings under the same spatial reference. It also writes the depth, volume, and radar reflection characteristics of deep diseases into the corresponding locations. Through multi-mode data association, it provides the connectivity judgment and risk level of surface cracks and deep seepage channels, realizing a three-dimensional visualization of hidden danger distribution output similar to CT scan, providing complete and reliable information for location, assessment, and repair decisions.
[0109] Example 1:
[0110] To verify the feasibility of this invention in practice, it was applied to an engineering scenario of intelligent inspection and hidden danger detection of dams. The typical characteristics of this scenario are that there may be visible defects such as surface cracks in the dam body, as well as deep hidden dangers such as piping, cavities, and seepage channels. Deep hidden dangers can only be identified by three-dimensional ground-penetrating radar imaging. Traditional methods often require professionals to spend a long time interpreting radar images and confirming them by drilling, which is inefficient and inconsistent. At the same time, the distribution of radar images of dams with different soil types varies significantly, and the accuracy of the model can drop sharply when it is migrated to a new area. In addition, the categories of defects are naturally uneven in real engineering projects, with many samples of common categories and few samples of niche categories, resulting in obvious category bias in the general detection network, which cannot meet the needs of "surface and deep" full-domain hidden danger location and risk expression.
[0111] In this scenario, the present invention uses a mobile inspection platform equipped with a 3D ground-penetrating radar to detect the dam body and simultaneously acquire the GPS coordinates of the dam corresponding to the acquisition location. At the same time, it acquires visual image data for surface defect labeling. The 3D radar imaging data output by the 3D ground-penetrating radar expresses the spatial scatterer distribution in the form of point cloud and the radar profile imaging results along the inspection path in the form of grayscale profile map. The radar imaging data is first subjected to wavelet threshold denoising to suppress electromagnetic interference and random noise. Then, the denoised radar imaging data and the dam GPS coordinates are coordinate registered and unified to a preset coordinate reference and preset data format to obtain a standardized radar feature map. On the visual image side, a large visual model identifies and labels the surface defects, outputting the labeling position, labeling length, and labeling width of the surface defects. Subsequently, based on the dam GPS coordinates, a correspondence between the visual image and the standardized radar feature map under the same acquisition location is established. The surface defect location labeling information is mapped to the spatial range of the standardized radar feature map and fused with it to form a standardized multi-mode feature map, which includes radar channel, i.e., reflection intensity data, thereby providing a unified input for subsequent models.
[0112] To address the problem of confusion between defects and background caused by poor adaptability of radar imaging feature extraction, this invention inputs a standardized multi-mode feature map into the radar feature enhancement attention module RFE-Attention. Reflection intensity data is extracted from the radar channel and feature decomposition is performed. The variance of the reflection values within the window is calculated using a 3x3 window sliding method on the reflection intensity data as a local scattering feature. Simultaneously, the phase difference of the reflected wave along the depth direction is calculated as a global layered feature. The local scattering feature is used to characterize local reflection anomalies caused by defects, while the global layered feature is used to characterize the layered structure of the dam body and suppress local deviations caused by structural errors. A dual-branch weight generator performs weight mapping on the two types of features respectively, obtaining local attention weights and layered attention weights, which are then multiplied to form the RFE-Attention weight map. The weight fusion unit multiplies this weight map pixel-by-pixel with the standardized multi-mode feature map and writes it into the enhanced feature map, thereby achieving input optimization of "defect feature enhancement and background and inter-layer noise suppression" at the data level.
[0113] To address the class bias problem caused by small sample sizes and class imbalance, this invention enhances the feature map input to the class-specific detection architecture CS-VIT. First, a shared feature extraction layer performs convolution and pooling on the enhanced feature map and compresses the feature channels to obtain a shared feature map. Then, class-specific training branches are set according to the number of disease categories. Each branch includes a VIT encoder, an existence judgment head, and a target box regression head. The VIT encoder divides the shared feature map into blocks to form a patch sequence and performs Transformer encoding to obtain class-specific features. The existence judgment head outputs the existence of the disease category and the detection confidence. The target box regression head outputs the regression results used to determine the upper left and lower right corner positions of the target box. The cross-branch fusion layer performs weighted voting fusion of the detection results of each branch and outputs the detection boxes and category labels for all disease categories. This allows niche categories to obtain more concentrated and stable feature learning and decision-making paths within their dedicated branches, reducing the suppressive effect of multi-class shared training on small sample categories.
[0114] To address the issue of imaging distribution drift caused by cross-regional environments, this invention performs cross-domain adaptive training and optimization on CS-VIT and forms a detection-update closed loop. 3D radar imaging data with disease category and target bounding box annotations are used as labeled radar data, while 3D radar imaging data without annotation information are used as unlabeled radar data, forming multiple regional datasets based on their geographical origin. The domain adaptation network migrates and aligns the imaging distribution of unlabeled radar data from different regions to output cross-domain enhanced data. The radar physical enhancement operator sets propagation difference parameters based on different soil conditions and simulates radar wave propagation differences, addressing the issue of strong reflection... The attenuation is adjusted and regionally specific noise is superimposed to generate regionally adaptive augmented samples. During training, dynamic weights are set for each category, which increase as the sample frequency decreases and as the detection confidence decreases. A constant term is introduced to avoid the weights being zero. Dynamic weighting alleviates category bias. At the same time, a phased model is used to generate detection results for unlabeled and cross-domain augmented data, and high-confidence results are selected as pseudo-labels. The pseudo-labels are mixed with labeled data for iterative training and updated. Finally, the optimized CS-VIT model is obtained, and the original multi-category disease detection results are updated. The updated multi-category disease detection results are output.
[0115] To quantify and verify the beneficial effects of this invention, an evaluation dataset covering multiple types of dam soil conditions was constructed without specifying a particular time or location. Comparative experiments were conducted. Scheme A (baseline) uses a multi-class shared detection model with conventional feature extraction and standard attention mechanisms without cross-domain adaptive training and optimization. Scheme B adds RFE-Attention to the baseline to enhance the standardized multi-modal feature map with radar features before detection. Scheme C replaces the detection model with a CS-VIT-type dedicated detection architecture to detect multiple types of defects. Scheme D uses both RFE-Attention feature enhancement and a CS-VIT-type dedicated detection architecture for detection, but without cross-domain adaptive training and optimization. Scheme E (this invention) further adds cross-domain adaptive training and optimization and a detection update closed loop to Scheme D, and outputs the updated multi-class defect detection results to generate a three-dimensional hazard distribution map of the dam body. This invention introduces RFE-Attention, CS-VIT, and cross-domain adaptive training and optimization under the same training budget. Specific comparative data are shown in Table 1.
[0116] Table 1. Comparison of Key Performance Indicators of Different Hazard Detection Schemes
[0117] index Option A Option B Option C Option D Option E Overall detection accuracy mean mAP (%) 61.8 69.5 71.2 76.8 83.6 Recall (%) 58.4 66.2 68.9 74.1 81.3 False negative rate (%) 23.7 17.5 16.2 12.8 7.9 False positive rate (%) 17.9 15.1 14.6 12.3 9.8 Average AP (%) for niche categories 18.6 24.3 39.7 46.9 57.8 Cross-domain accuracy decrease (percentage points) 31.4 26.8 24.9 18.6 6.7 Connectivity detection accuracy (%) 62.1 68.7 70.5 78.4 86.9 Automatic processing time per kilometer (min) 38 41 44 46 52
[0118] As shown in Table 1, the general model exhibits low overall accuracy and recall in radar imaging scenarios, with a significant decrease in accuracy after cross-domain transfer, particularly in the average detection capability of niche categories. Introducing RFE-Attention improves both overall accuracy and recall while reducing the false negative rate, indicating that integrating radar physical features and hierarchical information into attention directly improves "disease-background confusion" and "inter-layer noise interference." Introducing CS-VIT significantly increases the average AP for niche categories, demonstrating that class-specific branches and existence checks can alleviate class bias caused by class imbalance. When RFE-Attention and CS-VIT are combined, the model further improves in both overall and niche category metrics, indicating that input feature enhancement and class-specific detection are complementary in mechanism. Further introduction of cross-domain adaptive training and optimization, along with detection updates, significantly reduces the decline in cross-domain accuracy, suggesting that domain adaptation alignment and radar physical enhancement can narrow regional imaging distribution differences. Dynamic weighting and pseudo-label iteration can continuously improve model generalization under minimal labeling conditions, ultimately achieving more stable detection output without increasing manual interpretation workload.
[0119] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for detecting hidden dangers in the entire dam area based on radar image interpretation and visual large model, characterized in that, Includes the following steps: Data on the embankment was collected using a mobile inspection platform equipped with a 3D ground-penetrating radar, resulting in 3D radar imaging data, visual images, and the corresponding GPS coordinates of the embankment. The 3D radar imaging data is denoised and then registered with the GPS coordinates of the dam to obtain a standardized radar feature map. The location information of surface diseases is obtained based on visual images and spatially aligned with standardized radar feature maps to form standardized multi-mode feature maps. The standardized multi-mode feature map is input into the radar feature enhancement attention module RFE-Attention, and feature enhancement processing is performed through the feature decomposition unit, the dual-branch weight generator and the weight fusion unit to output the enhanced feature map. The enhanced feature map input class-specific detection architecture CS-VIT is used to extract features through a shared feature extraction layer, class-specific training branches and cross-branch fusion layers to obtain multi-class disease detection results; Cross-domain adaptive training and optimization were performed on CS-VIT, and the detection results of multiple diseases were updated based on the optimized CS-VIT to obtain the updated detection results of multiple diseases. The updated multi-category disease detection results are correlated with the dam's GPS coordinates and surface disease location marking information to generate a three-dimensional hazard distribution map of the dam body.
2. The method for detecting hidden dangers in the entire dam area based on radar image interpretation and visual large model according to claim 1, characterized in that, The acquisition of the three-dimensional radar imaging data, visual images, and corresponding GPS coordinates of the dam specifically includes: Based on the 3D radar data acquisition module, a mobile inspection platform equipped with 3D ground-penetrating radar is provided. It moves along the dam inspection path and collects data during the movement. During the acquisition process, visual images and the GPS coordinates of the dam corresponding to the acquisition location are acquired simultaneously. The three-dimensional ground-penetrating radar operates at a center frequency of 200MHz to 500MHz to conduct three-dimensional detection of the dam body within a depth range of 0m to 5m, obtaining three-dimensional radar imaging data. The three-dimensional radar imaging data is output as a point cloud to show the spatial scattering distribution of the dam body, and as a grayscale profile map to show the radar profile imaging results along the inspection path.
3. The method for detecting hidden dangers in the entire dam area based on radar image interpretation and visual large model according to claim 1, characterized in that, The standardization of radar feature maps is obtained specifically through: Wavelet thresholding is performed on the 3D radar imaging data to eliminate electromagnetic interference. Wavelet decomposition is performed and the coefficients obtained by decomposition are thresholded. Inverse wavelet reconstruction is performed on the thresholded coefficients, and the denoised 3D radar imaging data is output. The denoised 3D radar imaging data is subjected to coordinate registration processing to determine the correspondence between each acquisition location in the denoised 3D radar imaging data and the GPS coordinates of the dam, and the data is bound together and the bound radar data is output. The bound radar data is standardized by unifying it with a preset coordinate reference and a preset data format, and the standardized result is output as a standardized radar feature map.
4. The method for detecting hidden dangers in the entire dam area based on radar image interpretation and visual large model according to claim 1, characterized in that, The formation of the standardized multimodal feature map specifically includes: The surface disease is marked on the visual image. The marking includes the marking position, marking length and marking width of the surface disease, and the surface disease location marking information is output. Based on the GPS coordinates of the dam, establish the correspondence between visual images and standardized radar feature maps at the same acquisition location; Based on the correspondence, the surface disease location annotation information is mapped to the spatial range of the standardized radar feature map to obtain the mapped surface disease location annotation information, which is then fused with the standardized radar feature map to output a standardized multi-mode feature map.
5. The method for detecting hidden dangers in the entire dam area based on radar image interpretation and visual large model according to claim 1, characterized in that, The output of the enhanced feature map specifically includes The standardized multi-mode feature map is input into the radar feature enhancement attention module RFE-Attention, which includes a feature decomposition unit, a dual-branch weight generator and a weight fusion unit, and extracts radar channel, i.e., reflection intensity data, from the standardized multi-mode feature map as data to be processed. The feature decomposition unit performs feature decomposition on the reflection intensity data. A three-to-three window is used to perform sliding calculation on the reflection intensity data. The variance of the reflection value within the window is used as the local scattering feature, and the phase difference of the reflected wave along the depth direction is used as the global layering feature. The dual-branch weight generator performs weight mapping processing on the local scattering features and the global hierarchical features respectively. The local scattering features are mapped by the Sigmoid function to output local attention weights, and the global hierarchical features are mapped by the Softmax function to output hierarchical attention weights. The local attention weights and hierarchical attention weights are multiplied to obtain the RFE-Attention weight map. The weight fusion unit multiplies the RFE-Attention weight map pixel by pixel with the standardized multimodal feature map and outputs the enhanced feature map.
6. The method for detecting hidden dangers in the entire dam area based on radar image interpretation and visual large model according to claim 1, characterized in that, The results of the multi-category disease detection specifically include: The enhanced feature map is input into the class-specific detection architecture CS-VIT, which is a large visual model including a shared feature extraction layer, class-specific training branches set according to the number of disease categories, and a cross-branch fusion layer. The class-specific training branches include a VIT encoder, an existence judgment head, and a target box regression head. In the shared feature extraction layer, convolution and pooling are performed sequentially on the enhanced feature map. Through four layers of convolution and pooling operations, the feature channels are compressed into 256 channels to obtain the shared feature map. In each dedicated training branch, the shared feature map is divided into blocks according to the preset block division rules, forming a block sequence, so that the block sequence contains sixty-four blocks. In the VIT encoder, the tile sequence is encoded using a six-layer Transformer to obtain the class-specific features corresponding to the current disease category; In the existence judgment head, the class-specific features are processed sequentially with fully connected processing and Sigmoid activation processing to obtain the judgment result of whether the current disease category exists, and the detection confidence corresponding to the judgment result is obtained. In the target bounding box regression head, the class-specific features are processed four times in succession to obtain the regression results used to determine the upper left and lower right corner positions of the target bounding box, forming the branch detection results of the current disease category; In the cross-branch fusion layer, the branch detection results output by various dedicated training branches are subjected to weighted voting fusion processing. The results are weighted according to the detection confidence and the optimal result is selected. The detection boxes and category labels of all categories of diseases are output as multi-category disease detection results.
7. The method for detecting hidden dangers in the entire dam area based on radar image interpretation and visual large model according to claim 1, characterized in that, The updated multi-category disease detection results are obtained specifically including: Read the detection results of multiple types of diseases, and identify the three-dimensional radar imaging data with disease category and target box annotation information as labeled radar data, otherwise identify it as unlabeled radar data. Then divide the unlabeled radar data into different regional data sets according to the regional source to form the input data set for cross-domain adaptive training and optimization. Perform cross-domain data augmentation processing, call the domain adaptation network to migrate and align the radar imaging distribution of different regional datasets, and output the aligned unlabeled radar data as cross-domain augmentation data; The radar physical enhancement sample generation process is executed. The radar physical enhancement operator is called to enhance the labeled radar data and cross-domain enhancement data. The reflection intensity attenuation is adjusted in the propagation difference simulation process, and regional specific noise is superimposed to output regional adaptive enhancement samples. Perform dynamic weight loss training preparation processing, set training loss weights corresponding to the current disease class for each disease class, and introduce constant terms to avoid training loss weights being zero. Dynamic weight loss training is performed. During the training process, the sample frequency of each disease category is counted in real time, and the detection confidence of each disease category is counted in real time. The loss weights corresponding to each disease category are updated. The updated loss weights are used to weight the classification loss and regression loss of CS-VIT. At the same time, the model training is updated. The labeled radar data and the regional adaptive enhancement samples are used as training data input to obtain the CS-VIT model with phased optimization. Semi-supervised pseudo-label iterative training is performed. The CS-VIT model with phased optimization is used to detect cross-domain augmented data. The detection results are filtered according to the detection confidence. The detection results with a detection confidence of not less than 0.85 are selected as pseudo-labels. The labeled radar data and pseudo-label data are mixed to form an iterative training set. The CS-VIT model is iteratively trained and updated. The pseudo-labels are updated after each iteration of training to obtain the optimized CS-VIT model. The optimized CS-VIT model is used to update the detection results of multiple disease categories, and cross-branch fusion processing is performed to output the updated detection results of multiple disease categories.
8. The method for detecting hidden dangers in the entire dam area based on radar image interpretation and visual large model according to claim 1, characterized in that, The generation of the three-dimensional hazard distribution map of the dam body specifically includes: The updated multi-category disease detection results are processed by coordinate association. The spatial location corresponding to each disease detection result is bound to the GPS coordinates of the dam and written into the data record corresponding to the GPS coordinates of the dam, forming a set of disease detection records with coordinates. Simultaneously perform coordinate association processing on the surface disease location annotation information to form a set of surface disease annotation records with coordinates; Based on the coordinate-based disease detection records and coordinate-based surface disease location annotation information, a three-dimensional aggregation process is performed. The disease category, target box and detection confidence corresponding to the same dam GPS coordinates are aggregated. The surface disease annotation location, annotation length and annotation width corresponding to the same dam GPS coordinates are aggregated. The aggregated surface disease information and deep disease information are distributed and organized in three dimensions according to the spatial location of the dam to obtain a three-dimensional hidden danger information set. The three-dimensional hazard information set is processed using CT-style output. The annotation position, annotation length and annotation width of the surface disease are marked by visual image and written into the corresponding spatial position. The annotation depth, annotation volume and radar reflection characteristics of the deep disease are written into the corresponding spatial position. The surface cracks are judged to be connected to the deep seepage channels by multi-mode data association and written into the corresponding spatial position. The location, depth and risk level information are written into the corresponding spatial position to generate a three-dimensional hazard distribution map of the dam body.