Bridge foundation water seepage intelligent detection method and system

By enhancing the consistency between infrared and visible light images through adaptive thermal excitation and phase-locked sampling, and combining thermal and wet fingerprint extraction and segmentation networks, candidate seepage masks and uncertainty maps are generated. This solves the problem of insensitivity to weak signs in bridge foundation seepage detection and achieves efficient and stable localization of seepage areas in bridge foundations.

CN121473396AInactive Publication Date: 2026-02-06WANNIAN COUNTY PENGRUI CONSULTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511469844.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-02-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional bridge foundation seepage detection methods are not sensitive to weak seepage signs in underwater or humid environments, and the multimodal detection results are difficult to form a verifiable unified expression in three-dimensional space. They lack explicit expression and propagation of model prediction uncertainty and cannot effectively locate the apparent area of ​​the bridge foundation.

Method used

By enhancing the consistency between infrared and visible light images through adaptive thermal excitation and phase-locked sampling, and combining thermal and wet fingerprint extraction and segmentation networks, candidate water seepage masks and uncertainty maps are generated. Voxel-level energy modeling and consistency scoring are then performed to achieve unified indexing and weighting of cross-modal evidence.

Benefits of technology

It improves the accuracy and stability of bridge foundation seepage detection, generates verifiable three-dimensional seepage area location results, reduces data redundancy, and constructs a traceable closed loop for the entire process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121473396A_ABST
    Figure CN121473396A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of civil engineering structure detection, and particularly discloses a bridge foundation water seepage intelligent detection method and system. The method comprises the following steps: acquiring field multi-sensor data and carrying out calibration and synchronous processing to generate a region-of-interest map; carrying out trigger acquisition and alignment check according to the initial value sequence, and generating a structure detection quantity initial value sequence; performing drift estimation and compensation to generate a drift mask; extracting and segmenting the hot and wet fingerprints based on the drift mask, and generating candidate water seepage masks; mapping the candidate water seepage mask to a radar sector for sampling configuration, performing cross-modal consistency comparison with a structure detection quantity, and performing weighting and energy construction on a voxel set based on a region-of-interest map to generate a voxel energy map; and finally, outputting a detection result through optimization solution and space fitting. According to the method, through multi-modal data fusion and cross-modal consistency verification, the accuracy and reliability of water seepage detection are effectively improved, and false alarm and missing alarm are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of civil engineering structure detection technology, and in particular to a bridge foundation water seepage intelligent detection method and system. BACKGROUND

[0002] In the technical field of civil engineering structure detection technology, the bridge foundation is long-term in the underwater or humid, alternating shadow and significant illumination change environment, the seepage signs generated on the surface and near surface layer of the bridge foundation are often weak in amplitude, unstable in persistence, and obviously affected by temperature and light drift. The traditional method of image threshold segmentation relying on manual inspection or single sensor is not sensitive to weak seepage areas and is easily disturbed by reflected highlights, shadows, algae coverage and water surface disturbance; when only relying on ground penetrating radar echo to judge near surface layer anomalies, the positioning results and the apparent area of the bridge foundation are difficult to strictly correspond due to the influence of medium unevenness and attitude change. The existing multi-modal attempts are mostly concentrated on feature fusion and segmentation at the frame level or pixel level, lacking a voxel level constraint carrier for unified modeling in time series and spatial three dimensions, lacking explicit expression and propagation of model prediction uncertainty, and lacking a systematic method for consistent registration, weight allocation and time sequence punishment of image side candidate regions and radar side physical evidence under a unified index.

[0003] In engineering applications, infrared images and visible light images are affected by temperature drift and light drift, and insufficient or excessive compensation will introduce artifacts; ground penetrating radar echoes appear interlayer echo shape changes under rough external surfaces and complex geometries, which are difficult to correspond one by one with image partitions; structural detection quantities are mostly derived from local window apparent statistics, and if there is a lack of unified index and weight mechanism, the correspondence between them and image candidate regions, radar features is unstable. The above problems make it difficult for multi-modal results to form a unified expression that can be reviewed in three-dimensional space, and it is difficult to provide directly callable energy targets for subsequent boundary refinement and spatial positioning.

[0004] In view of the above problems, adaptive thermal excitation and phase-locked sampling are gradually introduced in the industry to enhance the thermal correlation of weak responses, and to improve the consistency of infrared and visible light images through time sequence alignment and environmental drift compensation; at the same time, a thermal and moisture fingerprint extraction and segmentation network is introduced to obtain candidate regions, and an uncertainty map is output as a measure of the confidence of the model output. However, the existing process often stays at the level of candidate water seepage masks, and has not yet built a mechanism to fuse candidate water seepage masks, uncertainty maps and evidence from the radar side and structural detection quantity side into the same three-dimensional index system, and to weight and energy register voxels in the time and space dimensions, lacking an engineering path to convert the consistency score sequence into voxel-level physical side weights and form a unified energy expression with image side weights. SUMMARY

[0005] To address the aforementioned technical problems, this invention provides an intelligent detection method for bridge foundation seepage, comprising:

[0006] The system acquires the raw response and calibration reference data of the infrared cameras, visible light cameras and ground seepage radar deployed on site, performs geometric calibration, radiometric calibration, unified time baseline establishment, excitation and sampling synchronization relationship calculation, field of view analysis and occlusion processing, and generates a region of interest map.

[0007] Based on the region of interest map, the system calls adaptive thermal excitation parameters and phase-locked sampling parameters to perform trigger configuration and coherent acquisition, calls unified time baseline and time sequence markers to perform time sequence alignment and spatial alignment, and references partition information and thermal excitation time period markers to perform local clipping and attitude verification to generate an initial value sequence of structural detection quantities.

[0008] Based on the initial value sequence of structural detection, temperature drift and illumination drift are estimated by calling the aligned image sequence and synchronous echo sequence; radiation compensation and illumination compensation are performed by partition / segment / response type; strong / weak masks are generated by frame-by-frame inspection at the partition level to mark drift areas and process noise modeling, and drift masks are generated.

[0009] Based on the drift mask, thermal and wet fingerprint extraction and local screening, segmentation model training and inference processing are performed to generate candidate water seepage masks and uncertainty maps.

[0010] Based on the candidate permeation mask, the pixel region is mapped to the sector and distance range by calling the joint calibration parameters to perform sector sampling and step size configuration, cross-modal comparison is performed with the initial value sequence of structural detection to perform consistency calculation, and voxel weighting and energy construction are performed using the region of interest map as the outer frame discrete voxel set to generate a voxel energy map.

[0011] Based on the voxel energy map, optimization and boundary refinement, spatial fitting and voxel clustering, statistical and parameter update processing are performed to generate a parameter write-back package.

[0012] Furthermore, the raw response and calibration reference data of the infrared cameras, visible light cameras, and ground seepage radar deployed on-site include:

[0013] The raw responses and calibration reference data of the infrared cameras, visible light cameras, and ground-penetrating radar deployed on-site specifically include calibration image sequences and their corresponding imaging parameters collected by the infrared and visible light cameras, which contain multiple angles, distances, and attitudes; echo reference sequences and their corresponding transmission and reception parameters collected by the ground-penetrating radar; spatial coordinates and feature point distribution data of the calibration reference components required for geometric calibration in the detection area; reference measurement data of reference targets with stable emission characteristics and reference materials with stable reflection characteristics required for radiation calibration under various temperature differences and various surface wet conditions; and the original registration data of the internal time stamp verification records of each device used to establish a unified time baseline, the delay interval measurement data from trigger to data availability, the acquisition start time, the acquisition end time, and the trigger sequence mark.

[0014] Furthermore, sector sampling and step size configuration include:

[0015] The pixel regions marked in the candidate seepage mask are mapped to the sectors and distance ranges corresponding to the seepage radar, forming a set of candidate sectors;

[0016] At the region level, the region of interest map is invoked, and regions marked as maskable are removed. For regions marked as risk observation areas, only candidate segments that persist across frames are retained.

[0017] Read the triggering and acquisition periods of the synchronization echo sequence at the time level to make the candidate sector set correspond one-to-one with the corresponding time period;

[0018] Based on the above three types of correspondence, a sampling strategy is configured for each candidate sector: when the overlap area between the candidate sector and the candidate seepage mask is large, dense sampling with a small step size is set; when the overlap area is small or located in the risk observation area, sparse sampling with a large step size is set, and the sampling reason is recorded simultaneously.

[0019] Furthermore, the process of performing cross-modal alignment and consistency calculation with the initial sequence of structural detection values ​​also includes:

[0020] After the sampling strategy is determined, the original records of the synchronous echo sequence are read along the distance and time directions. The changes in echo intensity, echo arrival position, interlayer continuity and reflection pattern are calculated. The results of multiple samplings within the same sector are denoised and merged to generate feature segments with time markers, sector markers and distance markers.

[0021] The feature fragment is bound to the spatial index of the candidate water seepage mask, so that each radar feature can be traced back to the candidate region and time segment from which it originated.

[0022] For regions that are removed or sparsely processed, the mask state and sampling step size are registered in the feature fragment;

[0023] All feature segments are organized according to sector order and time order to obtain local radar feature sequences and output them.

[0024] Furthermore, the process of performing cross-modal alignment and consistency calculation also includes:

[0025] Establish a one-to-one correspondence between each feature segment in the local radar feature sequence and the candidate region bound to it; at the time level, based on the previously registered time stamps, compare segments from the radar side and the image side within the same candidate region in the same time segment.

[0026] The apparent quantities related to boundary stability, texture change and local darkening in the initial value sequence of structural detection quantities are read and mapped to the local window of each candidate region. They are then compared item by item with the echo intensity change, arrival position change and interlayer continuity of the local radar feature sequence.

[0027] Furthermore, the process of performing cross-modal alignment and consistency calculation also includes:

[0028] For each comparison, a source identifier, region identifier, and time period identifier are attached to the recording end. Based on the comparison results, a hierarchical registration of comparison entries is generated within the spatial grid of the candidate region. First, the direction and magnitude of change on the image side and the radar side are matched within the frame, and then the continuity and repetition are matched between frames. Entries marked as risk observation areas by the region of interest are weighted down during registration. Entries registered as low confidence by the candidate water seepage mask are given source hints during registration. Finally, the above entries are summarized in each candidate region according to spatial grid and time segment to form a scoring record with weights and identifiers. The record is organized into a consistency scoring sequence according to the regional order and time order and output.

[0029] Furthermore, the process of performing voxel weighting and energy construction using the region of interest map as the outer frame discrete voxel set also includes:

[0030] At the spatial level, using the region of interest map as the outer frame, the spatial range covered by the candidate region is discretized into a set of voxels conforming to the registration specifications of this system, and an index field from the image side and the radar side is established for each voxel. At the temporal level, using the time stamp registered in the previous order as a reference, the candidate state and consistency of the same voxel in adjacent frames are recorded as time entries. Then, weighting is carried out at the voxel level: image-side weighting is based on the candidate water penetration mask, and the candidate state of the voxel in the corresponding frame is registered as the image-side weight base value, and the base value is attenuated or maintained by the uncertainty map; physical-side weighting is based on the consistency score sequence, and the score records corresponding to the voxel in the same region are aggregated by time entries and registered as the physical-side weight base value.

[0031] Furthermore, when a voxel is located in a shieldable region or a risk observation area of ​​the region of interest map, the weight base values ​​of the image side and the physical side are uniformly weakened or delayed in effect according to the registration rules of that region, and the reasons for weakening or delaying are recorded.

[0032] Furthermore, the process of generating a voxel energy map also includes:

[0033] After completing the image-side and physical-side weighting, the weights of voxels on the time entries are checked for continuity. Time periods in which the same voxel is continuously candidate across frames and is continuously supported by the physical side are marked as continuous entries, while voxels that appear only in individual frames are marked as isolated entries. Different time penalty rules are registered for continuous entries and isolated entries at the recording end. After completing the above weighting, a unified energy registration value is formed for each voxel based on its image-side weight base value, physical-side weight base value, and time entry penalty record. The voxel set is then organized into a continuous voxel energy map according to the region and time order and output.

[0034] Furthermore, a bridge foundation seepage intelligent detection system, applied to any of the methods described above, includes:

[0035] The multimodal data acquisition and calibration module is used to acquire the raw response and calibration reference data of infrared cameras, visible light cameras and ground penetration radar, and to perform geometric calibration, radiometric calibration and unified time baseline establishment to generate joint calibration parameters.

[0036] The synchronous acquisition and control module is used to calculate the synchronization relationship between excitation and sampling based on joint calibration parameters to generate adaptive thermal excitation parameters and phase-locked sampling parameters, and to generate synchronous image sequences and synchronous echo sequences based on the region of interest map for trigger configuration and coherent acquisition.

[0037] The data alignment and preprocessing module is used to perform temporal and spatial alignment based on the synchronized image sequence and joint calibration parameters to generate an aligned image sequence, and to perform local cropping and pose verification based on the aligned image sequence to generate an initial value sequence of structural detection quantities.

[0038] The drift estimation and compensation module is used to estimate temperature drift and illumination drift based on the initial value sequence of structural detection quantities to generate environmental drift parameters, perform radiation compensation and illumination compensation based on the environmental drift parameters to generate a compensated image sequence, and perform drift region labeling and noise modeling based on the compensated image sequence to generate a drift mask.

[0039] The feature extraction and model module is used to extract and locally filter thermal and wet fingerprint features based on the compensated image sequence and drift mask. Based on the thermal and wet fingerprint features, the module trains the segmentation model to generate a pre-trained segmentation model. Based on the pre-trained segmentation model, the module performs inference to generate candidate seepage masks and uncertainty maps.

[0040] The consistency verification module is used to generate local radar feature sequences by sector sampling and step size configuration based on candidate seepage masks, to generate consistency score sequences by consistency calculation based on local radar feature sequences, and to generate voxel energy maps by voxel weighting and energy construction based on consistency score sequences.

[0041] The optimization and localization module is used to perform optimization and boundary refinement based on the voxel energy map to generate a refined mask sequence, and to perform spatial fitting and voxel clustering based on the refined mask sequence to generate a set of seepage areas and a spatial location sequence.

[0042] The parameter management module is used to perform statistical analysis and parameter updates based on the set of seepage areas and spatial location sequences, generating parameter write-back packages.

[0043] The key innovations of this invention include:

[0044] (1) The three-source fusion voxel weighting mechanism of candidate seepage mask-uncertainty map-consistency score sequence unifies image confidence, physical evidence and time item penalty into voxel energy map to serve subsequent optimization solution and boundary refinement;

[0045] (2) A coherent acquisition strategy driven by adaptive thermal excitation parameters and phase-locked sampling parameters, combined with the region of interest map, is used to construct stable time segments and region clipping entry points for weak signs;

[0046] (3) Radiation compensation and illumination compensation driven by environmental drift parameters, and residual unstable regions are labeled by drift mask to constrain training samples and inference masking rules;

[0047] (4) Sector sampling and step size configuration are combined with the initial value sequence of structural detection to form a consistency scoring sequence, so that the evidence from the radar side and the structure side can be registered in an engineering manner under the spatial index and time index of the candidate area;

[0048] (5) The parameter write-back package is designed for multi-stage linkage updates of data acquisition, compensation, verification and energy modeling. It transforms the statistical conclusions of the voxel energy map into the parameter inputs for the next round of processes, and constructs a traceable closed loop for the entire process.

[0049] The following are its main beneficial effects:

[0050] (1) Effectiveness of archiving and acquisition. By establishing a unified time baseline and spatial mapping through joint calibration parameters, a region of interest map is generated. With the help of adaptive thermal excitation parameters and phase-locked sampling parameters, triggering and coherent acquisition are organized to form a synchronous echo sequence and an aligned image sequence. The initial value sequence of structural detection quantities is also registered to reduce data redundancy caused by invalid observations and to build a stable entry point for subsequent temporal alignment and region clipping.

[0051] (2) Practicality of environmental disturbance suppression and sample constraints. Based on the aligned image sequence, synchronous echo sequence and initial value sequence of structural detection, the environmental drift parameters are estimated. Radiation compensation and illumination compensation are performed on infrared and visible light to generate compensated image sequences. The residual unstable regions are marked at the frame level and the partition level by drift mask to limit the interference of abnormal samples on subsequent feature construction and model training.

[0052] (3) Targeted fingerprint construction and candidate generation. Based on the compensated image sequence and adaptive thermal excitation, thermal and wet fingerprint features are extracted. Local screening is performed by combining drift mask and outputting candidate water seepage masks and uncertainty map. The uncertainty map explicitly describes the confidence of the model output and provides a quantifiable image-side weight basis for subsequent cross-modal verification and voxel weighting.

[0053] (4) Engineering verifiability of cross-modal collaborative verification. Sector sampling and step size configuration are performed on the candidate seepage mask and synchronous echo sequence to generate local radar feature sequence; consistency score sequence is calculated by combining the initial value sequence of structural detection quantity, and the evidence support status of radar side and structure side is registered under the spatial index and time index of candidate region to form physical side criteria that correspond one-to-one with the image side results.

[0054] (5) Core Value of Voxel-Level Energy Modeling. Based on the most innovative steps outlined in the claims, voxel weighting and energy construction are performed using candidate permeation masks, uncertainty maps, and consistency scoring sequences to generate a voxel energy map: a set of discrete voxels is defined within the spatial bounding box of the region of interest map, and candidate states and consistency are registered within time entries; candidate permeation masks are converted into image-side weight base values, which are then attenuated or maintained by the uncertainty map; the consistency scoring sequence is aggregated into physical-side weight base values; uniform weakening or delayed activation is applied to maskable areas and risk observation areas, and the reasons are recorded; time-based penalty records are generated for cross-frame persistent and isolated performance; unified energy registration values ​​are formed at the voxel level, organized into a voxel energy map, maintaining frame-level binding with candidate permeation masks and region-level binding with consistency scoring sequences. This voxel energy map directly carries the target description for subsequent optimization and boundary refinement, supports the spatiotemporal consistency constraints of spatial fitting and voxel clustering, and retains the source index required for tracing, meeting the needs of engineering verification.

[0055] (6) Implementation path of localization and closed-loop update. Based on the voxel energy map and candidate seepage masks, optimization and boundary refinement are carried out to output a refined mask sequence. Spatial fitting and voxel clustering are performed on the refined mask sequence to form a set of seepage areas and spatial location sequence. Combined with the consistency score sequence, statistical source evidence and time item distribution are used to output parameter write-back package, which provides archived update suggestions for adaptive thermal excitation parameters, phase-locked sampling parameters, environmental drift parameters, sector sampling and step size configuration, voxel weighting and time penalty settings, so that the acquisition, compensation, segmentation, verification and localization form a closed-loop flow of data and parameters under a unified index. Attached Figure Description

[0056] Figure 1 A flowchart illustrating an intelligent detection method for bridge foundation seepage provided in an embodiment of this application;

[0057] Figure 2 This is a structural block diagram of an intelligent bridge foundation seepage detection system provided in an embodiment of this application. Detailed Implementation

[0058] Example 1: Refer to Figure 1 This is a flowchart illustrating an intelligent detection method for bridge foundation seepage provided in an embodiment of the present invention. The process may include at least steps S100-S600:

[0059] S100: Acquire the raw response and calibration reference data of the infrared camera, visible light camera and ground seepage radar deployed on site, perform geometric calibration, radiometric calibration, unified time baseline establishment, excitation and sampling synchronization relationship calculation, field of view analysis and occlusion processing, and generate region of interest map;

[0060] S200: Based on the region of interest map, the adaptive thermal excitation parameters and phase-locked sampling parameters are called to perform trigger configuration and coherent acquisition; the unified time baseline and time sequence markers are called to perform time sequence alignment and spatial alignment; the partition information and thermal excitation time period markers are referenced to perform local clipping and attitude verification processing, and the initial value sequence of structural detection quantities is generated.

[0061] S300: Based on the initial value sequence of structural detection quantities, call the aligned image sequence and synchronous echo sequence to estimate temperature drift and illumination drift, perform radiation compensation and illumination compensation according to partition / segment / response type, and generate strong / weak masks by frame-by-frame inspection at the partition level to perform drift region labeling and noise modeling processing, and generate drift masks.

[0062] S400: Based on the drift mask, perform hot and wet fingerprint extraction and local screening, segmentation model training and inference processing, and generate candidate water seepage masks and uncertainty maps.

[0063] S500: Based on the candidate permeation mask, the joint calibration parameters are called to map the pixel region to the sector and distance range for sector sampling and step size configuration; cross-modal comparison is performed with the initial value sequence of structural detection quantities for consistency calculation; and voxel weighting and energy construction are performed using the region of interest map as the outer frame discrete voxel set to generate a voxel energy map.

[0064] S600 performs optimization and boundary refinement, spatial fitting and voxel clustering, statistical and parameter update processing based on voxel energy maps, and generates parameter write-back packages.

[0065] Step S100 includes at least steps S110-S130:

[0066] S110. Acquire the raw response and calibration reference data of the infrared camera, visible light camera and ground seepage radar deployed on site, perform geometric calibration, radiometric calibration and establish a unified time baseline, and generate joint calibration parameters.

[0067] The raw response and calibration reference data of the infrared cameras, visible light cameras, and ground seepage radar deployed on-site specifically include the following: calibration image sequences and their corresponding imaging parameters collected by the infrared and visible light cameras, including multiple angles, distances, and attitudes; echo reference sequences and their corresponding transmission and reception parameters collected by the ground seepage radar; spatial coordinates and feature point distribution data of the calibration reference components required for geometric calibration in the detection area; reference measurement data of reference targets with stable emission characteristics and reference materials with stable reflection characteristics required for radiation calibration under various temperature differences and various surface wettability conditions; and the original registration data of the internal time stamp verification records of each device used to establish a unified time baseline, the delay interval measurement data from trigger to data availability, the acquisition start time, the acquisition end time, and the trigger sequence mark.

[0068] Specifically, the acquisition and archiving process for subsequent multimodal processing is completed first. During operation, infrared cameras, visible light cameras, and ground-penetration radar are deployed in the bridge foundation detection area, forming stable supports according to the detection points and equipment installation posture to ensure the fixed relationship between the equipment and the bridge foundation surface. Subsequently, geometric calibration and radiometric calibration are carried out sequentially, and a unified time baseline is established. During the geometric calibration process, calibration reference components set in the detection area guide the infrared and visible light cameras to acquire calibration images containing multiple angles, distances, and orientations. Through multi-viewpoint observation, matching data of corresponding image coordinates and spatial coordinates are formed, and the echo reference sequence of the ground-penetration radar is collected simultaneously to form a correspondence between image reference coordinates and radar reference coordinates. Through the above correspondence, geometric calibration parameters are formed. The geometric calibration parameters are used to define the relative positional relationship, relative orientation relationship, and imaging scale relationship between the infrared and visible light cameras, and to define the mapping relationship between image coordinates and bridge foundation structure coordinates, as well as the mapping relationship between image coordinates and ground-penetration radar sector coordinates. Next, radiometric calibration is performed. Using a reference target with stable emission characteristics and a reference material with stable reflection characteristics, response data from the infrared and visible light cameras are collected under various temperature differences and surface wetness conditions. A correspondence between changes in incident energy and changes in image grayscale is established to obtain radiometric calibration parameters. During this process, the stable and drift intervals of the camera response are recorded through repeated measurements, and non-uniform response locations are marked to ensure that consistent intervals of the radiometric response can be used in subsequent processing. To ensure the temporal availability of multimodal data, a unified time baseline is further established: a unified triggering link and time stamping rules are configured between the infrared camera, visible light camera, and ground penetration radar. The time stamps within each device are verified, the delay interval from triggering to data availability for each device is measured, and the acquisition start time, acquisition end time, and trigger sequence marker are recorded in a unified format to complete the registration of the unified time baseline. Through the above steps, the geometric calibration parameters, radiometric calibration parameters, and unified time baseline are structurally integrated to obtain joint calibration parameters. At this point, the input to S110 is the raw response and calibration reference data of the infrared camera, visible light camera and ground penetration radar after on-site deployment, and the output is the joint calibration parameters. These joint calibration parameters will be used in S120 to calculate the excitation and sampling synchronization relationship, and will also serve as the alignment basis in the temporal alignment and spatial alignment in S200, and as a prerequisite for radiation compensation and spatial mapping in the compensation processing in S300.

[0069] S120. Based on the joint calibration parameters, calculate the synchronization relationship between excitation and sampling to generate adaptive thermal excitation parameters and phase-locked sampling parameters;

[0070] In step S120, using the joint calibration parameters as input, the synchronization relationship between excitation and sampling is calculated to obtain adaptive thermal excitation parameters and phase-locked sampling parameters. Specifically, firstly, the geometric calibration parameters and radiometric calibration parameters from the joint calibration parameters are called to determine the available response ranges of the infrared and visible light cameras, identify the stable range of image response and the range sensitive to temperature changes, and map the sector coverage range and range resolution range of the ground-penetrating radar accordingly. Based on the above mapping, the working rhythm planning of the thermal excitation device is completed, binding the start time, duration, and interval of thermal excitation to a unified time baseline, so that the opening, holding, and closing of thermal excitation form a one-to-one correspondence with the exposure rhythm of the infrared camera, the framing rhythm of the visible light camera, and the transmission and reception rhythm of the ground-penetrating radar, thereby determining the adaptive thermal excitation parameters. Subsequently, based on the unified time baseline in the joint calibration parameters, the sampling start point, sampling rhythm, and sampling window for coherent sampling are determined. The continuous view from the infrared camera is decomposed into acquisition units containing coherent frame groups, and the continuous view from the visible light camera is decomposed into acquisition units corresponding to coherent frame groups. The transmission and reception sequences of the ground-penetrating radar are time-bound with the coherent frame groups to form phase-locked sampling parameters. To ensure subsequent multimodal coupling and use, the reference relationship with the joint calibration parameters is further recorded in the adaptive thermal excitation parameters and phase-locked sampling parameters. The image acquisition sequence and echo acquisition sequence corresponding to each thermal excitation sequence are clearly defined, as are the trigger marker and time marker corresponding to each coherent frame group, thereby achieving synchronization of thermal excitation, image acquisition, and echo acquisition under a unified time baseline. Through this process, the input of S120 is the joint calibration parameters, and the output is the adaptive thermal excitation parameters and the phase-locked sampling parameters. The adaptive thermal excitation parameters and the phase-locked sampling parameters will be used in S200 for trigger configuration and coherent acquisition, and in S400 for weighting and filtering constraints during thermal and wet fingerprint extraction. At the same time, their temporal binding relationship will be referenced by S300 to ensure that the data correspondence before and after compensation is not destroyed, and finally referenced by S500 and S600 to maintain a consistent index between candidate water seepage masks, local radar feature sequences and voxel energy maps.

[0071] S130. Based on the joint calibration parameters, perform field of view analysis and occlusion detection to generate a region of interest map;

[0072] In S130, using joint calibration parameters as input, field-of-view analysis and occlusion detection are performed for subsequent acquisition and processing, generating a region of interest map. Specifically, firstly, the geometric calibration parameters in the joint calibration parameters are called. Based on the installation positions and attitudes of the infrared and visible light cameras, their respective field-of-view coverage areas and boundary lines are calculated. Then, according to the mapping relationship between image coordinates and bridge foundation structure coordinates in the joint calibration parameters, the field-of-view coverage areas are projected onto the bridge foundation structure coordinate system to obtain the field-of-view coverage areas of the infrared and visible light cameras. At the same time, the sector coverage area of ​​the ground infiltration radar is mapped to the same structural coordinate system, forming an overlapping area between the image field of view and the echo sector. Subsequently, combined with the radiometric calibration parameters in the joint calibration parameters, stable response areas and non-uniform response areas are identified within the field-of-view coverage area. Reflected highlight areas, long-term shadow areas, and water surface reflection areas that may introduce abnormal responses are marked to form the zoning results of stable observation areas and risk observation areas. Based on the above zoning results, occlusion detection is performed: by comparing the outlines and boundary lines of bridge foundation components in the calibration images, areas obscured by temporary supports, construction equipment, or water surface undulations are identified, and this occlusion information is registered as shieldable areas. Integrating field of view coverage, response zoning, and occlusion information, a region of interest (ROI) map is generated at both image coordinate and structural coordinate levels according to the marking format defined by the unified time baseline. This ROI map uses a grid or segmented approach to label the image areas of infrared and visible light cameras with the echo sectors of the ground seepage radar, clearly defining the spatial locations of stable observation areas, risk observation areas, and shieldable areas. An index relationship is established with the joint calibration parameters to ensure that subsequent processes can directly call upon this map for region cropping, sector selection, and shielding configuration. In this process, the input to S130 is the joint calibration parameters, and the output is the region of interest (ROI) map. The ROI map is invoked by the trigger configuration and coherent acquisition in S200 for region cropping and occlusion masking of the synchronized image sequence and synchronized echo sequence. It is also used by S220 for cropping and mapping preservation of the aligned image sequence, and in S500 as the sector gating basis for candidate water penetration masks and the spatial prior for voxel weighting. To ensure data closure, S130 retains reference fields for the adaptive thermal excitation parameters and phase-locked sampling parameters when generating the ROI map, ensuring that the region definition for each time period is consistent with each coherent frame group, thereby avoiding temporal and spatial mismatches in subsequent stages.

[0073] Step S200 includes at least steps S210-S230:

[0074] S210. Based on the region of interest map, perform trigger configuration and coherent acquisition to generate a synchronous image sequence and a synchronous echo sequence;

[0075] S210 outlines the implementation steps for trigger configuration and coherent acquisition. This step takes adaptive thermal excitation parameters and phase-locked sampling parameters as inputs, and combines them with the unified time baseline and geometric calibration parameters in the joint calibration parameters, as well as the region of interest map, to perform configuration and acquisition. Specifically, firstly, the adaptive thermal excitation parameters are called to set the start time, duration, and interval of the thermal excitation device, and the markers for each time period are bound to the unified time baseline in the joint calibration parameters, so that the on / off, hold, and off of the thermal excitation are recorded under the same time reference as the exposure rhythms of the infrared and visible light cameras, as well as the transmission and reception rhythms of the infiltration radar. Subsequently, the phase-locked sampling parameters are called to group and configure the acquisition rhythms of the infrared and visible light cameras, dividing continuous acquisition into coherent frame groups, and writing the start and end markers of each coherent frame group into the time marker sequence; at the same time, the transmission and reception of the infiltration radar are gated to form echo records corresponding one-to-one with the image acquisition within the time window covered by the coherent frame group. To reduce unnecessary data entering subsequent processes, the region of interest (ROI) map is further invoked to mask and register the exposure areas of the infrared and visible light cameras and the sector scans of the ground-permeable radar. Without altering the subsequent cropping and alignment responsibilities, these maskable areas are marked during the acquisition phase so that corresponding labels can be added to the records, ensuring direct reference for subsequent alignment and cropping. After completing the above configuration, the acquisition process is executed according to the adaptive thermal excitation parameters and phase-locked sampling parameters: the infrared and visible light cameras synchronously capture images within coherent frame groups and record time markers, field-of-view markers, and device attitude markers; the ground-permeable radar completes transmission and reception within the time window corresponding to the coherent frame group and records echo markers and sector markers. At the end of the acquisition, the continuous frame groups from the infrared and visible light cameras are organized into a synchronized image sequence, and the continuous echo records from the ground-permeable radar are organized into a synchronized echo sequence, maintaining the index relationship between the two types of sequences and the joint calibration parameters, as well as the regional relationship with the ROI map. Therefore, the output of S210 is a synchronized image sequence and a synchronized echo sequence. The synchronized image sequence is used as the input of S220 for temporal and spatial alignment. The synchronized echo sequence is referenced in subsequent steps with time stamps and sector stamps and forms a one-to-one correspondence with the image results in subsequent coordination steps.

[0076] S220. Based on the synchronized image sequence and joint calibration parameters, perform temporal alignment and spatial alignment to generate an aligned image sequence;

[0077] S220 is the implementation step for temporal and spatial alignment. This step takes the synchronized image sequence and joint calibration parameters as input, and completes the unified mapping and consistent registration of data from the infrared and visible light cameras while ensuring that the time stamps and region markers generated in S210 are not destroyed. Specifically, firstly, the unified time baseline in the joint calibration parameters is called, and the start and end markers and intra-frame markers of each coherent frame group in the synchronized image sequence are read. The time stamps generated by different devices are converted into a time series under a unified recording rule, and the order between coherent frame groups, the order within coherent frame groups, and the interval between adjacent frames are restored on this time series, so that the synchronized image sequence forms non-overlapping continuous segments on the time axis. Subsequently, the geometric calibration parameters in the joint calibration parameters are invoked to perform a unified transformation on the imaging relationship between the infrared and visible light cameras. The two types of images are mapped to the projection plane corresponding to the bridge foundation structure coordinates. During the mapping process, the partition markers of the region of interest (ROI) map are used as mask constraints. For partitions registered as shieldable areas, only the region markers are retained without spatial geometric solutions. Complete geometric transformation solutions and boundary records are performed for stable and risky observation areas, respectively, thus establishing pixel-level and region-level correspondences between the infrared and visible light images at the spatial level. To facilitate subsequent steps, coherent frame group markers and acquisition attitude markers from S210 are further maintained in the alignment record, enabling accurate retrieval of coherent frame groups on a unified time baseline and clearly distinguishing the mapping results under different observation attitudes. After completing temporal and spatial alignment, the continuous images processed by the unified time baseline and geometric calibration parameters are organized into an aligned image sequence. Simultaneously, the partition correspondence between each frame and the ROI map, as well as the index relationship with the joint calibration parameters, are retained in this aligned image sequence. Therefore, the output of S220 is an aligned image sequence, which is used as the input of S230 for local cropping and pose verification. At the same time, the time markers and region markers in the aligned image sequence are directly referenced in the drift estimation and environment compensation of S300 to ensure the consistency of the inter-frame correspondence and spatial mapping relationship before and after compensation.

[0078] S230. Based on the aligned image sequence, perform local cropping and pose verification to generate an initial value sequence of structure detection quantities;

[0079] S230 is the implementation step for local cropping, attitude verification, and generating an initial sequence of structural detection values. This step takes the aligned image sequence as input and uses the partition information in the region of interest map to perform local processing and attitude consistency checks. Specifically, firstly, the aligned image sequence is partitioned and located according to the region of interest map. Local cropping is performed in stable observation areas, and boundary-preserving local cropping is performed in risky observation areas. In partitions registered as maskable areas, only the region marker is recorded and no cropped fragment is output, thus forming a set of local images that correspond one-to-one with the partitions. The time marker, partition marker, and attitude marker of each local image are fully inherited. Subsequently, attitude verification is performed on each local image: the geometric calibration parameters in the joint calibration parameters and the device attitude marker are called to compare the boundary position and structure line of the same partition in adjacent coherent frame groups. When the angle difference, scale difference, or boundary offset exceeds the registered threshold, the local image of that partition is marked as a segment that needs to be verified and the verification result is recorded. For local images that pass the verification, their partition boundaries, texture boundaries, and temporal continuity are retained for subsequent statistics. Based on this, and combined with the thermal excitation time period markers registered in S210, local images in different thermal excitation on and off segments are aggregated into regions. An initial value sequence of structure detection quantities is generated based on directly calculable appearance indicators such as image grayscale changes, texture changes, and boundary stability. This initial value sequence of structure detection quantities is registered simultaneously with the partition markers, time markers, and attitude markers to ensure that retrieval can be performed by time, partition, and attitude in subsequent stages. To maintain continuity with subsequent processes, an inter-frame correspondence index with the aligned image sequence is established during the generation of the initial value sequence of structure detection quantities, enabling each initial value record to be traced back to its source frame and source partition. After completing the above processing, S230 outputs an initial sequence of structure detection values ​​and submits this output, along with the time stamp and region stamp of the aligned image sequence, to subsequent steps. This allows S300 to perform drift estimation and environmental compensation using the aligned image sequence and the initial sequence of structure detection values ​​as input. It also enables S400 to perform thermal and wet fingerprint extraction and candidate segmentation using local segments with clear partitioning and pose consistency after compensation. Furthermore, it allows S500 to perform sector gating and region weighting based on partition stamps during the collaborative verification and voxel weighting stages. Understandably, S230 simultaneously retains a time reference field for the synchronous echo sequence at the recording end, allowing subsequent cross-modal processes to establish a consistent index between image partitions and echo sectors.

[0080] Step S300 includes at least steps S310-S330:

[0081] S310. Based on the initial value sequence of structural detection quantities, temperature drift and illumination drift are estimated to generate environmental drift parameters.

[0082] The S310 takes the aligned image sequence, synchronous echo sequence, and initial value sequence of structural detection quantities as input. Under the index relationship consistent with the joint calibration parameters, adaptive thermal excitation parameters, phase-locked sampling parameters, and region of interest map, it performs temperature drift and illumination drift estimation to obtain environmental drift parameters. Specifically, the aligned image sequence is first fragmented in chronological order, so that consecutive frames affected by the same thermal excitation segment and the same sampling window within the same time period form candidate groups. The stable observation area marked in the region of interest map is used as the reference area. The grayscale response, boundary response, and texture response of the reference area in adjacent candidate groups are compared to establish the response difference trajectory of the same area in different time segments. At the same time, the sector records of the same bridge position are retrieved in the synchronous echo sequence corresponding to the reference area. The echo amplitude and arrival pattern caused by medium changes are compared to mark structural disturbances that may not be related to changes in thermal environment or illumination, so as to avoid misusing structural disturbances as drift drivers. Subsequently, the response difference trajectories on the image side and the sector records on the echo side were compared in parallel. Difference spectra were established for image segments in different regions of thermal excitation activation and deactivation. Boundary stability and texture stability indices in the initial value sequence of structural detection quantities were used as screening criteria to filter out abnormal segments formed during periods of abrupt changes in component geometry or sensing attitude, resulting in a set of effective segments suitable for drift estimation. Based on this, for temperature drift estimation, the slow shift trend of the response level was statistically analyzed along the time axis for stable observation areas of infrared images; for illumination drift estimation, the gradual changes in brightness, shadow, and reflection in visible light images under indirect and semi-direct illumination conditions were stratified and summarized; and sector records in the corresponding synchronous echo sequence without significant medium changes were used as exclusion markers to distinguish response changes caused by ambient temperature and external illuminance from those caused by the state of the measured object. Furthermore, the effective fragment sets are aggregated along the bridge location and partition dimensions to form partition descriptions of temperature drift and illumination drift. The correspondence between these descriptions and the structural coordinates is recorded using a coordinate mapping consistent with the joint calibration parameters. Finally, the partition descriptions and temporal segment relationships are uniformly registered to form environmental drift parameters. These environmental drift parameters, indexed by temporal segment, partition location, and response type, maintain a one-to-one correspondence with the aligned image sequence and serve as the sole compensation basis for S320. They also provide a referenceable drift description for feature selection and template gain adjustment in S400 and an exclusion marker for drift sources in consistency calculation in S500.

[0083] S320. Based on environmental drift parameters, perform radiation compensation and illumination compensation to generate a compensated image sequence;

[0084] S320 takes environmental drift parameters and aligned image sequences as input, and performs radiometric compensation and illumination compensation on infrared and visible light images respectively according to the index relationship of partition, time and response type to obtain a compensated image sequence. In implementation, the time segment definition of environmental drift parameters is first read to ensure that the compensation operation is consistent with the aforementioned candidate grouping and to avoid cross-segment mixing. Then, based on the temperature drift description in the environmental drift parameters, the infrared image is segmented and corrected in the stable observation area and the risk observation area respectively: in the stable observation area, the slowly shifting response level is translated and compressed in chronological order, and the correction amount is added to the image frame of the segment in the form of partitioned records; in the risk observation area, small-amplitude correction is performed with boundary preservation as the criterion, and the correction traces are marked separately for subsequent screening. Simultaneously, based on the illumination drift description in the environmental drift parameters, brightness equalization, local shadow suppression, and conservative reflection suppression are performed on the visible light image in areas of indirect, semi-direct, and potentially specular reflection, respectively. For areas with occlusion or reflected highlights, only mask markings are recorded without detailed repair to maintain the recognizability of abnormal areas in subsequent models. To ensure multimodal consistency, the compensation results for infrared and visible light images are spatially registered consistently according to joint calibration parameters, and the inconsistency of partition boundaries before and after compensation is checked. If partition boundary drift occurs, the partition boundaries are reset based on the region of interest map to ensure that the compensated image still corresponds one-to-one with the predetermined partitions. During the compensation process, adaptive thermal excitation parameters and phase-locked sampling parameters are combined to distinguish the different effects of thermal excitation enabled and disabled segments on the image response: for thermal excitation enabled segments, subtle differences reflecting thermal response are preferentially retained, and only overall offsets related to drift are corrected; for thermal excitation disabled segments, a more balanced compensation strategy is implemented to facilitate stable stitching between subsequent segments. After the above processing, the infrared and visible light images, compensated by partitioning, segmenting, and response type, are re-merged into a continuous compensated image sequence. Each frame retains a traceability marker indicating its compensation source, clarifying whether the compensation originates from temperature drift or illumination drift, and the partitioning strategy employed. This compensated image sequence serves as the sole input to S330 for drift region labeling and noise modeling. Simultaneously, in S400, this sequence serves as a direct input for thermal and humidity fingerprint extraction and candidate segmentation, and the retained traceability markers provide a reference for sample selection and loss weight allocation in subsequent models. Furthermore, in S500, this sequence will be used as a region alignment reference in consistency calculations in conjunction with the synchronous echo sequence, and in S600, it will be used for comparison during the boundary refinement stage.

[0085] S330. Based on the compensated image sequence, perform drift region annotation and noise modeling to generate a drift mask;

[0086] The S330 takes the compensated image sequence as input, marks the drift regions for possible residual environmental influences, and performs noise modeling to generate drift masks. First, at the partition level, the compensated image sequence is inspected frame by frame: within stable observation areas, the grayscale texture and boundary consistency of adjacent frames are compared chronologically, and segments showing short-term jumps or local abnormal enhancements are registered as candidate drift regions; within risk observation areas, combined with local occlusion and reflection markers, only abnormal regions that persist across frames are retained, filtering out isolated segments that suddenly appear in a single frame. Then, at the intra-frame level, refinement is performed: using the difference mapping before and after compensation within the same frame as a reference, the boundary neighborhood and texture neighborhood are checked for signs of insufficient or excessive compensation, and cross-validation is performed with multiple frames within the same partition to eliminate false judgments of boundary displacement caused by pose changes. After completing the above partitioning and intra-frame filtering, a set of candidate drift regions is formed. Noise modeling is performed based on this set: For the infrared image channel, candidate regions are merged according to response amplitude and temporal persistence to characterize the slight temperature disturbances that may still exist after compensation; for the visible light image channel, candidate regions are merged according to brightness gradient and local contrast to characterize the shadow and reflection residues that may still exist after compensation; and the spatial and temporal coverage of these noise components are registered at the partition level so that they can be used as the basis for weight reduction or sample screening in subsequent processing. After noise modeling is completed, a drift mask is generated based on the candidate drift region set and the registered noise components: For stable observation areas, candidate regions that appear continuously across frames and are consistent over multiple frames are marked as strong masks, and candidate regions that appear intermittently are marked as weak masks; for risky observation areas, weak masks are marked only within the range that does not overlap with previous occlusion or reflection marks, and unresolved marks are retained for review in subsequent stages. To ensure consistency with the entire workflow, the drift mask and the compensated image sequence are kept frame-level aligned, and a partition index is established with the region of interest map, so that subsequent steps can selectively enable or disable mask constraints within specific partitions. The drift mask is directly provided to the thermal and wet fingerprint extraction and segmentation model training in S400 for local filtering and sample weight setting. It also serves as an external prior in the voxel weighting stage of S500 for weight reduction and as a constraint hint for unstable regions in the boundary refinement stage of S600. To ensure data closure, S330, while outputting the drift mask, appends the partition information, segment information, and noise component information used to generate the mask in a unified format to the frame-level annotation of the compensated image sequence, facilitating cross-step traceability and verification.

[0087] Step S400 includes at least steps S410-S430:

[0088] S410. Based on the compensated image sequence and combined with the drift mask, perform thermal and wet fingerprint extraction and local screening to generate thermal and wet fingerprint features;

[0089] The S410 takes a compensated image sequence and adaptive thermal excitation parameters as input, and combines drift masking to perform thermal and wet fingerprint extraction and local screening, outputting thermal and wet fingerprint features. Specifically, firstly, the compensated image sequence is divided into time segments according to the adaptive thermal excitation parameters, so that consecutive frames in the same excitation segment and the same sampling window are merged into consistent segments, and the time stamp and partitioning mark of the segments are kept consistent with the existing index. Then, within each consistent segment, candidate region sets are established for stable observation areas and risk observation areas according to the region of interest map. In the stable observation area, texture changes, boundary changes, and temperature response changes are directly read, while in the risk observation area, only sub-regions not marked as unstable are retained as a restriction by the drift mask. On this basis, the temporal change trajectory related to temperature response is extracted from the infrared channel in the compensated image sequence, and the appearance change trajectory related to wet marks, darkening, and edge diffusion is extracted from the visible light channel at the same spatial location. The two types of trajectories are aligned by time stamp within the segment and then stitched point by point to form a set of fingerprint segments with a unified spatial index and a unified temporal index. Furthermore, to avoid disturbances caused by occlusion and reflection, the drift mask is used as a masking condition on the same spatial index. Extraction is stopped for areas marked as strong masks, and only the change trajectory that exists consistently across frames is retained for areas marked as weak masks. The masking and retention states are then appended to the fingerprint segment. After segment-level processing, all fingerprint segments are merged at the partition level: first, the continuity of the change trajectory of adjacent segments is checked in the stable observation area to ensure that the trajectory is traceable in time and located in space; then, cross-checking is performed in the risk observation area to remove short-term anomalies that overlap with the drift mask. After merging, the selected temporal changes, appearance changes, and boundary changes are encapsulated into thermal and wet fingerprint features with a unified structure, and registered together with time markers, partition markers, and masking markers to form the output of S410. The thermal and wet fingerprint features serve as the sole input of S420, and also as the direct basis for sample selection and weight setting in subsequent steps, maintaining a one-to-one correspondence with the compensation image sequence, facilitating backtracking to the original partition and original time segment when needed.

[0090] S420. Based on the thermal and wet fingerprint features, a segmentation model is trained to generate a pre-trained segmentation model;

[0091] The S420 uses thermal and wet fingerprint features and a drift mask as input to train a segmentation model, resulting in a pre-trained segmentation model. Specifically, the thermal and wet fingerprint features are first organized into samples. Fingerprint segments from different time segments and partitions are uniformly arranged according to spatial indices, ensuring that each sample group contains fingerprints related to temperature response, wetness appearance, and boundary diffusion. These samples are then linked to corresponding frames in the compensation image sequence for traceability. Subsequently, based on the drift mask, sample selection and weighting are performed during the sample organization stage: spatial locations with strong mask coverage are not included in training; spatial locations with weak mask coverage but consistent across frames have reduced weights while preserving their temporal continuity; and locations not covered by the mask maintain normal weights and undergo sample equalization within the same partition to avoid partition bias during training. Furthermore, during the training preparation stage, segment labels from the adaptive thermal excitation parameters are introduced into the sample organization table, enabling the model to distinguish between excitation-on and excitation-off segments when reading samples, and to learn the fingerprint differences between the two types of segments under the same spatial index. During training, samples are iterated according to a unified partition and time index order. This allows the model to prioritize learning continuous trajectories and stable boundary regions in stable observation areas, and to prioritize learning persistent trajectories filtered by drift masks in risky observation areas. After each iteration, the masking markers and weights of the samples are written back to the training record to ensure auditable sample usage. To maintain consistency with subsequent inference stages, the model's input organization, partition index, and time index are permanently registered after training. This clarifies the model's frame-level alignment when reading compensated image sequences, its spatial alignment when reading thermal and wet fingerprint features, and its masking priority when reading drift masks. After completing the above process, the pre-trained segmentation model is output. This output is directly available for use by the S430, and its training record and sample organization table are retained within the same index system. This allows for frame-by-frame verification of the candidate region generation process when needed, and provides a clear source indication during subsequent collaborative verification with synchronous echo sequences.

[0092] S430. Based on the pre-trained segmentation model, inference is performed to generate candidate seepage masks and uncertainty maps;

[0093] S430 takes the compensated image sequence as input and feeds it into the pre-trained segmentation model to generate candidate water penetration masks and uncertainty maps. Specifically, it first reads the input organization and indexing method registered in S420 of the pre-trained segmentation model. Under this method, the compensated image sequence is loaded in groups according to partitions and time segments, so that each group of frames to be processed strictly corresponds to its partition label and time label when read. Then, in the inference stage, each group of frames to be processed and its spatial index are submitted to the pre-trained segmentation model simultaneously. The model internally organizes the infrared and visible light channels according to the trajectory agreed upon during training and performs collaborative parsing. It also calls the drift mask for real-time masking according to the masking priority recorded in the sample organization table, so that the positions marked as strong masks do not participate in candidate generation, and the positions marked as weak masks only participate in the weighted processing of candidate confidence. After obtaining the candidate responses for each frame, inter-frame consistency checks are performed according to the temporal order within the same frame group. Positions of continuous responses across frames under the same spatial index are integrated, and positions appearing only in a single frame are registered as low-confidence candidates. This registration status is recorded for later use in the collaborative verification phase. At the spatial level, candidate responses are uniformly mapped based on the geometric relationships in the joint calibration parameters, ensuring that responses from different devices fall into consistent positions within the same structural coordinate system, guaranteeing a one-to-one correspondence with the sector records of subsequent synchronous echo sequences. After completing the temporal and spatial consistency integration, the candidate responses for each frame are output as candidate permeation masks in binary or multi-level form. Simultaneously, an uncertainty map is output based on the confidence estimation and shielding participation within the model. The candidate permeation mask records the spatial range of the candidate region and the inter-frame continuity, while the uncertainty map records the confidence changes and shielding effects of the same spatial index within the same frame group. To ensure closed-loop operation, S430 binds the candidate water seepage mask, uncertainty map, and frame-level and partition indices of the compensated image sequence at the output, and retains references to thermal and wet fingerprint features, enabling subsequent steps to trace the specific fingerprint source. In the data flow direction, the candidate water seepage mask and uncertainty map serve as direct inputs to S500 for sector sampling, step size configuration, and consistency calculation. Simultaneously, in S600, they participate in localization and aggregation as initial constraints for optimization and boundary refinement.

[0094] Step 500 includes at least steps S510-S530:

[0095] S510. Based on the candidate water seepage mask, sector sampling and step size configuration are performed to generate local radar feature sequences.

[0096] In a method and system for intelligent detection of seepage in bridge foundations, the S510 takes candidate seepage masks and synchronous echo sequences as inputs. Under an index relationship consistent with joint calibration parameters, region of interest maps, and previously registered time stamps, it performs sector sampling and step size configuration to obtain local radar feature sequences. Specifically, firstly, at the spatial level, the joint calibration parameters are invoked to map the pixel regions marked in the candidate seepage masks to the sectors and distance ranges corresponding to the seepage radar, forming a set of candidate sectors. Secondly, at the regional level, the region of interest map is invoked to remove areas marked as maskable, and only candidate segments that persist across frames are retained for areas marked as risk observation areas. Finally, at the temporal level, the triggering and acquisition periods of the synchronous echo sequences are read to ensure a one-to-one correspondence between the set of candidate sectors and the corresponding time periods. Based on the above three types of correspondence, a sampling strategy is configured for each candidate sector: when the overlap area between the candidate sector and the candidate seepage mask is large, dense sampling with a small step size is set; when the overlap area is small or located in a risk observation area, sparse sampling with a large step size is set, and the sampling reason is recorded simultaneously. After the sampling strategy is determined, the original records of the synchronous echo sequence are read along the range and time directions, and directly measurable attributes such as echo intensity change, echo arrival position change, inter-layer continuity, and reflection morphology are calculated. Multiple sampling results within the same sector are denoised and merged to generate feature segments with time markers, sector markers, and range markers. To ensure cross-modal consistency, this feature segment is bound to the spatial index of the candidate seepage mask, so that each radar feature can be traced back to its source candidate region and time segment; for regions that are eliminated or sparsely processed, the mask state and sampling step size are registered in the feature segment. After the above processing is completed, all feature segments are organized in sector order and time order to obtain local radar feature sequences and output them. The local radar feature sequences serve as the only input to S520 and are also cited as a source of evidence on the physical side in the subsequent voxel weighting stage.

[0097] S520: Based on local radar feature sequences, perform consistency calculations to generate consistency score sequences;

[0098] The S520 takes the local radar feature sequence and the initial value sequence of structural detection quantities as input. Under the premise of maintaining consistency with the spatial and temporal indices of the candidate seepage mask, it performs consistency calculations to obtain a consistency score sequence. Specifically, firstly, at the spatial level, based on joint calibration parameters, a one-to-one correspondence is established between each feature segment in the local radar feature sequence and its associated candidate region. At the temporal level, based on previously registered time stamps, segments from the radar side and the image side within the same candidate region are compared within the same time segment. Subsequently, apparent quantities related to boundary stability, texture changes, and local darkening in the initial value sequence of structural detection quantities are read and mapped to the local window of each candidate region. These are then compared item by item with the echo intensity changes, arrival position changes, and interlayer continuity of the local radar feature sequence. In periods with attitude verification records, local windows that have passed verification are prioritized to avoid errors introduced by attitude changes. To ensure the verifiability of the comparisons, each comparison is accompanied by three types of identifiers on the recording end: first, a source identifier, indicating the image and radar segments referenced by the comparison; second, a region identifier, indicating the candidate region and region of interest where the comparison is located; and third, a time period identifier, indicating whether the comparison is in a thermal excitation on or off segment. Based on the above item-by-item comparison results, hierarchical registration of comparison entries is generated within the spatial grid of the candidate region. First, the direction and magnitude of change on the image and radar sides are matched within the frame, and then the persistence and repetition are matched between frames. Entries marked as risk observation areas by the region of interest are weighted weakly during registration; entries registered as low-confidence by the candidate water seepage mask are given source hints during registration. Finally, the above entries are summarized by spatial grid and time segment within each candidate region to form a weighted and labeled scoring record, organized into a consistency scoring sequence according to regional and temporal order and output. The consistency scoring sequence serves as one of the inputs to S530 and participates in weighting as the main source of physical consistency in subsequent voxel energy construction.

[0099] S530. Based on the consistency scoring sequence, perform voxel weighting and energy construction to generate a voxel energy map;

[0100] The S530 takes candidate water penetration masks, uncertainty maps, and consistency scoring sequences as inputs, performs voxel weighting and energy construction, and generates a voxel energy map. Specifically, firstly, at the spatial level, using the region of interest map as the outer frame, the spatial range covered by the candidate region is discretized into a set of voxels conforming to the system's registration specifications, and an index field from the image side and the radar side is established for each voxel; at the temporal level, using the previously registered time stamp as a reference, the candidate state and consistency of the same voxel in adjacent frames are recorded as time entries. Subsequently, weighting is carried out at the voxel level: image-side weighting is based on candidate water penetration masks, registering the candidate state of the voxel in the corresponding frame as the image-side weight base value, and attenuating or maintaining this base value using the uncertainty map; physical-side weighting is based on the consistency scoring sequence, aggregating the score records corresponding to the voxel in the same region according to time entries, and registering them as the physical-side weight base value; when a voxel is located in a maskable area or risk observation area of ​​the region of interest map, the weight base values ​​of the image side and the physical side are uniformly weakened or delayed according to the registration rules of that area, and the reasons for weakening or delaying are recorded. After weighting on both the image and physical sides, a continuity check is performed on the weights of voxels in time entries. Voxels that are consistently candidate across frames and continuously supported by the physical side are marked as persistent entries, while those appearing only in a few frames are marked as isolated entries. Different time penalty rules are registered for persistent and isolated entries at the recording end for direct reference in subsequent optimization and solution stages. To ensure consistency in subsequent spatial inference, spatial guidance consistent with the joint calibration parameters is restored on the voxel set, allowing the sources from the image and radar sides to be traced voxel-by-voxel when needed. After the above weighting is completed, a unified energy registration value is formed for each voxel based on its image-side weight base value, physical-side weight base value, and time entry penalty record. This value is then organized into a continuous voxel energy map on the voxel set according to region and time order and output. The voxel energy map is frame-level bound to the candidate water penetration mask and region-level bound to the consistency scoring sequence, serving as direct input for subsequent voxel localization optimization and boundary refinement. It can also be traced back to the local radar feature sequence and the initial value sequence of structural detection quantities when needed.

[0101] Understandably, steps S510 to S530 form a strict closed-loop link between input and output: S510 obtains the local radar feature sequence from the candidate seepage mask and synchronous echo sequence; S520 obtains the consistency score sequence from the local radar feature sequence and the initial value sequence of structural detection quantities; S530 generates a voxel energy map based on the candidate seepage mask, uncertainty map, and consistency score sequence. In summary, this step, from sector sampling and step size configuration, through cross-modal consistency calculation, to voxel weighting and energy construction, ensures that image-side candidates, physical-side evidence, and time entry constraints are structurally registered under the same indexing system, providing directly callable and traceable energy expressions and regional constraints for subsequent voxel localization optimization and parameter write-back.

[0102] Step S600 includes at least steps S610-S630:

[0103] S610. Based on the voxel energy map, perform optimization and boundary refinement to generate a refinement mask sequence;

[0104] The S610 takes a voxel energy map and candidate permeation masks as input. While maintaining index consistency with the joint calibration parameters and the region of interest (ROI) map, it performs optimization and boundary refinement, outputting a refined mask sequence. Specifically, firstly, at the spatial level, candidate permeation masks are used as initial constraints to determine the voxel set of the region to be optimized. Then, based on the energy records registered in the voxel energy map according to region and time sequence, continuous and isolated entries of the same voxel within adjacent frames are recovered as a continuity constraint in the temporal dimension. Subsequently, at the boundary level, pixel boundaries and voxel boundaries within the initial constraints are consistent: in stable observation areas, boundary smoothing and adjacency consistency constraints are prioritized to eliminate scattered edges; in locations marked as risky observation areas, the ROI map is used to preserve the original boundary orientation, and only jagged details that exist stably across frames are subjected to limited boundary straightening. The range and time period of this straightening process are recorded in the boundary revision record. Furthermore, at the energy level, a global solution is performed on the region to be optimized according to the voxel records of the voxel energy map: for voxels registered as persistent entries, their continuity influence is increased along the time segment, making the candidate states of that voxel in adjacent frames tend to be consistent; for voxels registered as isolated entries, their participation weight is weakened according to the penalty record, and their expansion is restricted spatially by adjacency consistency constraints. To ensure the correspondence of multimodal sources, the frame-level binding between the candidate seepage mask and the voxel energy map is maintained throughout the solution process, so that each boundary revision can be traced back to the corresponding time mark and partition mark. After the global solution is completed, a refined binary or multi-level mask is generated in each frame, and consistency verification is performed between frames in temporal order: regions that are persistent across frames retain their complete boundaries, and regions that appear only in a few frames and do not satisfy temporal continuity are eliminated or shrunk. Finally, the results of boundary refinement and time verification are organized into a refined mask sequence by frame and output. The refined mask sequence maintains the same index as the voxel energy map, candidate permeation mask and region of interest map, and is used for spatial fitting and voxel clustering calls of S620, and also serves as the starting basis for subsequent closed-loop statistics and parameter writing back of S600.

[0105] S620. Based on the refined mask sequence, spatial fitting and voxel clustering are performed to generate a set of seepage areas and a spatial location sequence.

[0106] The S620 takes a refined mask sequence as input and, under the constraint of the spatial mapping relationship with the joint calibration parameters, sequentially performs spatial fitting and voxel clustering to output a set of seepage areas and a spatial location sequence. Specifically, for each frame of the refined mask sequence, the connected mask regions are first decomposed in the image coordinates to obtain multiple candidate region groups; then, the joint calibration parameters are used to map each group to the bridge foundation structure coordinates, so that the boundary on the image side and the component outline on the structure side are compared and recorded in the same coordinate system. Subsequently, spatial fitting is performed: geometric fitting is performed on the boundary point set of each candidate region group to recover the main extension direction, boundary orientation, and local curvature changes of the region, and the assembly relationship with the bridge foundation components is checked in the mapped structural coordinates; for parts that coincide or approximately coincide with the component boundary, local adjustments are made on the premise of preserving the component boundary; for parts that have significant offset from the component surface, range rollback and boundary closure are performed without changing the time order of the refined mask sequence, and the above adjustments are recorded item by item in the fitting record. After spatial fitting is completed, fitted regions near the same structural location in consecutive frames are merged into the same voxel cluster based on proximity, similar boundaries, and temporal adjacency, thus implementing voxel clustering. During voxel clustering, regions registered as maskable are skipped directly using the region of interest map, while regions registered as risk observation areas are only merged if they persist across frames to prevent short-term segments from being mistakenly included. After voxel clustering is completed, the coverage, boundary orientation, and duration across frames are calculated for each voxel cluster, and its position index, range index, and time index are registered in the structural coordinates to form a spatial position sequence consistent with the temporal order. Simultaneously, the same voxel cluster is grouped into the same set entry, forming the seepage region set for this step. The above seepage region set and spatial position sequence maintain a bidirectional referencing relationship with the refined mask sequence: any set entry can be traced back to its source frame and source partition, and any mask segment within any frame can also be traced back to its set and corresponding position index. Therefore, the set of seepage areas and spatial location sequence output by S620 will be directly used as input to S630 for statistics and parameter updates, and will be cross-referenced with the consistency information registered in S500 outside the process.

[0107] S630. Based on the set of seepage areas and spatial location sequence, perform statistics and parameter updates to generate a parameter write-back package;

[0108] S630 takes the infiltration region set, spatial location sequence, and consistency score sequence as input. While ensuring consistency with the refined mask sequence, voxel energy map, and candidate infiltration mask indexes, it performs statistical analysis and parameter updates, generating a parameter write-back package. Specifically, at the region level, the infiltration region set and spatial location sequence are merged line by line to restore the start and end segments in time, the spatial coverage, and the position index in structural coordinates for each set entry. The corresponding record in the consistency score sequence is then retrieved using this index as the key, constructing a cross-modal comparison list. Subsequently, at the time level, the persistence, discontinuity, and boundary stability of regions are statistically analyzed in segment order. These statistical items are then paired with the region support of the consistency score sequence item by item: when a segment contains a high-support record, the segment is marked as a physically verified segment; when a low-support record or a record with weakened weights exists, the segment is marked as a segment to be verified, its source entry is referenced, and the verification index is retained for subsequent use. After completing cross-modal pairing, the parameter update phase begins: based on the temporal and positional distribution of the verified segments, update suggestion entries are generated for both the acquisition and compensation sides. The entries for the acquisition side point to the segment settings and acquisition window settings for adaptive thermal excitation parameters and phase-locked sampling parameters, while the entries for the compensation side point to the partition thresholds and segment divisions for environmental drift parameters. Simultaneously, based on the concentrated locations of the segments to be verified and low-support records, rhythm suggestions for sector sampling and step size configuration, as well as revision suggestions for voxel weighting ratio settings and time penalty settings, are generated. To ensure the write-back objects are directly accessible within the system, the aforementioned suggested items are organized into a parameter write-back package in a unified format. Each item within the package is labeled with the corresponding step and parameter, aligning adaptive thermal excitation parameters and phase-locked sampling parameters to acquisition-side steps, environmental drift parameters to compensation-side steps, sector sampling and step size configuration to collaborative verification-side steps, and voxel weighting and time penalty settings to energy modeling-side steps. Simultaneously, the write-back package retains a source index generated from the permeation region set, spatial location sequence, and consistency score sequence, ensuring that any updated item can be traced back to its statistical basis and cross-modal pairing results. After completing the statistical and parameter updates, the parameter write-back package is output, declaring its binding relationships with joint calibration parameters, region of interest maps, and various time markers, allowing for direct reading and execution by the acquisition, compensation, segmentation, verification, and positioning stages in the next workflow.

[0109] Example 2: Figure 2 A structural block diagram of a bridge foundation seepage intelligent detection system according to an embodiment of the present invention is shown. Figure 2 As shown, the structure may include:

[0110] The multimodal data acquisition and calibration module 01 is used to acquire raw response and calibration reference data from infrared cameras, visible light cameras, and ground-penetrating radar, and to perform geometric calibration, radiometric calibration, and the establishment of a unified time baseline to generate joint calibration parameters. Specifically, it receives raw response and calibration reference data from infrared cameras, visible light cameras, and ground-penetrating radar deployed on-site. Under the constraints of the configured detection points and equipment installation attitude, it completes the acquisition of multi-viewpoint observation calibration images, echo reference sequence acquisition, and multi-angle spatial coordinate matching to form geometric calibration parameters. Using reference materials with stable emission and reflection characteristics, it completes radiometric response data acquisition and repeated measurements under various temperature differences and surface wetness conditions to form radiometric calibration parameters. Through a unified trigger link and time stamping rules, it completes the time stamp verification and delay interval determination of each device to establish a unified time baseline. The geometric calibration parameters, radiometric calibration parameters, and unified time baseline are integrated into joint calibration parameters and transmitted to the synchronous acquisition control module for acquisition control. At the same time, the original registration information of the calibration reference data is retained for subsequent parameter traceability.

[0111] The synchronous acquisition control module 02 is used to calculate the synchronization relationship between excitation and sampling based on joint calibration parameters to generate adaptive thermal excitation parameters and phase-locked sampling parameters, and to generate synchronous image sequences and synchronous echo sequences based on the region of interest map (ROI). Specifically, it receives joint calibration parameters from the multimodal data acquisition and calibration module, calls geometric calibration parameters and radiometric calibration parameters to determine the available response range of the infrared and visible light cameras and map the coverage of the infiltration radar sector, completes the working rhythm planning of the thermal excitation device and the binding of the coherent sampling window according to a unified time baseline, and forms adaptive thermal excitation parameters and phase-locked sampling parameters; it completes the shielding registration of the exposure area and sector scanning in combination with the RIO map, executes the acquisition process according to the adaptive thermal excitation parameters and phase-locked sampling parameters, and forms synchronous image sequences and synchronous echo sequences; it transmits the synchronous image sequences and synchronous echo sequences to the data alignment and preprocessing module as alignment input, and registers time stamps and device attitude stamps in the acquisition record for subsequent modules to call.

[0112] The data alignment and preprocessing module 03 is used to perform temporal and spatial alignment based on the synchronized image sequence and joint calibration parameters to generate an aligned image sequence, and to perform local cropping and attitude verification based on the aligned image sequence to generate an initial value sequence of structural detection quantities. Specifically, it receives the synchronized image sequence and joint calibration parameters from the synchronized acquisition control module, calls the unified time baseline to complete the coherent frame group time stamp conversion and continuous segment recovery, calls the geometric calibration parameters to complete the projection mapping from infrared and visible light images to the bridge foundation structure coordinates, and uses the partition markers of the region of interest map as mask constraints to complete the geometric transformation solution of the stable observation area and the risk observation area to form an aligned image sequence; performs local cropping operation based on the partition information of the region of interest map, combines the device attitude markers to complete the verification and comparison of boundary positions and structural lines, and completes region aggregation and apparent index calculation based on the thermal excitation time period markers to generate an initial value sequence of structural detection quantities; the initial value sequence of structural detection quantities is passed to the drift estimation and compensation module as drift estimation input, while retaining the partition correspondence and time stamp of the aligned image sequence for subsequent compensation processing.

[0113] The drift estimation and compensation module 04 is used to estimate temperature drift and illumination drift based on the initial value sequence of structural detection quantities to generate environmental drift parameters, perform radiation compensation and illumination compensation based on the environmental drift parameters to generate a compensated image sequence, and perform drift region labeling and noise modeling based on the compensated image sequence to generate a drift mask. Specifically, the system receives the initial sequence of structural detection values ​​from the data alignment and preprocessing module. It then combines the aligned image sequence with the synchronous echo sequence to complete time-segmented processing and establish the response difference trajectory of the reference region. Structural disturbances are eliminated through echo sector record comparison, resulting in partitioned descriptions of temperature drift and illumination drift, which are registered as environmental drift parameters. Based on the partitioned and time-segmented definitions of the environmental drift parameters, segmented correction and brightness equalization operations are performed on the infrared and visible light images respectively. Compensation processing is completed by combining the differences between thermal excitation on and off segments, generating a compensated image sequence. On the basis of the compensated image sequence, frame-by-frame inspection and intra-frame refinement verification are performed. Strong and weak masks are generated by merging noise components, forming a drift mask. The drift mask is then passed to the feature extraction and model module as a filtering constraint, while the traceability markers of the compensated image sequence are retained for subsequent feature extraction.

[0114] Feature extraction and model module 05 is used to extract and locally filter hot and wet fingerprints based on compensated image sequences and drift masks to generate hot and wet fingerprint features. Based on the hot and wet fingerprint features, a segmentation model is trained to generate a pre-trained segmentation model. Based on the pre-trained segmentation model, inference is performed to generate candidate seepage masks and uncertainty maps. Specifically, the system receives compensated image sequences and drift masks from the drift estimation and compensation module, divides time segments based on adaptive thermal excitation parameters, establishes candidate region sets in stable and risky observation areas, and splices fingerprint segments by extracting the temporal change trajectory of the infrared channel and the appearance change trajectory of the visible light channel. It then generates thermal and wet fingerprint features by combining the drift mask shielding conditions. Based on these features, it organizes samples and sets weights, iteratively trains the segmentation model using partitioning and time indexing, and generates a pre-trained segmentation model. The compensated image sequences are grouped by partition and time segment and loaded into the pre-trained segmentation model. Candidate water seepage masks and uncertainty maps are generated through instantaneous masking and inter-frame consistency checks. These candidate water seepage masks and uncertainty maps are then passed to the consistency verification module for sector sampling, while the spatiotemporal index of the thermal and wet fingerprint features is retained for subsequent collaborative verification.

[0115] The consistency verification module 06 is used to generate local radar feature sequences by sampling sectors and configuring step sizes based on candidate permeation masks, calculate consistency scores based on these local radar feature sequences, and generate voxel energy maps by weighting voxels and constructing energy based on the consistency score sequences. Specifically, it receives candidate permeation masks from the feature extraction and modeling module, calls joint calibration parameters to complete the mapping from pixel regions to radar sectors, performs candidate sector elimination and retention processing in conjunction with the region of interest map, and generates local radar feature sequences by calculating echo intensity changes and inter-layer continuity. The local radar feature sequences are compared across modes with the initial value sequence of structural detection quantities, and the consistency score sequence is generated by registering source identifiers, region identifiers, and time period identifiers. Using the region of interest map as the outer frame of the discrete voxel set, the energy registration values ​​are calculated by weighting the image-side weight base values ​​and the physical-side weight base values, generating a voxel energy map. The voxel energy map is passed to the optimization and localization module as input for optimization, while retaining the region binding relationship of the consistency score sequence for subsequent voxel clustering.

[0116] The optimization and localization module 07 is used to perform optimization and boundary refinement based on the voxel energy map to generate a refined mask sequence, and then perform spatial fitting and voxel clustering based on the refined mask sequence to generate a set of seepage regions and a spatial location sequence. Specifically, it receives the voxel energy map from the consistency verification module, uses candidate seepage masks as initial constraints to complete the processing of continuous and isolated entries, and generates a refined mask sequence through boundary smoothing and adjacency consistency constraints; for each frame of the refined mask sequence, it performs region decomposition and geometric fitting, and performs voxel clustering based on the principles of distance proximity and boundary similarity to generate a set of seepage regions and a spatial location sequence; the set of seepage regions and the spatial location sequence are passed to the parameter management module for statistical input, while the temporal order of the refined mask sequence is retained for subsequent parameter updates.

[0117] The parameter management module 08 is used to perform statistical analysis and parameter updates based on the seepage area set and spatial location sequence to generate parameter write-back packets. Specifically, it receives the seepage area set and spatial location sequence from the optimization and positioning module, completes cross-modal pairing through time start and end segment statistics and boundary stability analysis, generates update suggestion entries for the acquisition side and compensation side based on the distribution of verified and unverified segments, and forms a parameter write-back packet. The parameter write-back packet is returned to the multimodal data acquisition and calibration module and the synchronous acquisition control module for parameter reinjection and link refresh, and completes the version update of the system strategy table.

Claims

1. A smart detection method for bridge foundation seepage, characterized in that, include: The system acquires the raw response and calibration reference data of the infrared cameras, visible light cameras and ground seepage radar deployed on site, performs geometric calibration, radiometric calibration, unified time baseline establishment, excitation and sampling synchronization relationship calculation, field of view analysis and occlusion processing, and generates a region of interest map. Based on the region of interest map, the system calls adaptive thermal excitation parameters and phase-locked sampling parameters to perform trigger configuration and coherent acquisition, calls unified time baseline and time sequence markers to perform time sequence alignment and spatial alignment, and references partition information and thermal excitation time period markers to perform local clipping and attitude verification to generate an initial value sequence of structural detection quantities. Based on the initial value sequence of structural detection, temperature drift and illumination drift are estimated by calling the aligned image sequence and synchronous echo sequence; radiation compensation and illumination compensation are performed by partition / segment / response type; strong / weak masks are generated by frame-by-frame inspection at the partition level to mark drift areas and process noise modeling, and drift masks are generated. Based on the drift mask, thermal and wet fingerprint extraction and local screening, segmentation model training and inference processing are performed to generate candidate water seepage masks and uncertainty maps. Based on the candidate permeation mask, the pixel region is mapped to the sector and distance range by calling the joint calibration parameters to perform sector sampling and step size configuration, cross-modal comparison is performed with the initial value sequence of structural detection to perform consistency calculation, and voxel weighting and energy construction are performed using the region of interest map as the outer frame discrete voxel set to generate a voxel energy map. Based on the voxel energy map, optimization and boundary refinement, spatial fitting and voxel clustering, statistical and parameter update processing are performed to generate a parameter write-back package.

2. The method according to claim 1, characterized in that, The raw response and calibration reference data of the infrared cameras, visible light cameras, and ground seepage radar deployed on-site include: The raw responses and calibration reference data of the infrared cameras, visible light cameras, and ground-penetrating radar deployed on-site specifically include calibration image sequences and their corresponding imaging parameters collected by the infrared and visible light cameras, which contain multiple angles, distances, and attitudes; echo reference sequences and their corresponding transmission and reception parameters collected by the ground-penetrating radar; spatial coordinates and feature point distribution data of the calibration reference components required for geometric calibration in the detection area; reference measurement data of reference targets with stable emission characteristics and reference materials with stable reflection characteristics required for radiation calibration under various temperature differences and various surface wet conditions; and the original registration data of the internal time stamp verification records of each device used to establish a unified time baseline, the delay interval measurement data from trigger to data availability, the acquisition start time, the acquisition end time, and the trigger sequence mark.

3. The method according to claim 1, characterized in that, Sector sampling and step size configuration include: Mapping to the sectors and ranges corresponding to the infiltration radar forms a set of candidate sectors; At the region level, the region of interest map is invoked, and regions marked as maskable are removed. For regions marked as risk observation areas, only candidate segments that persist across frames are retained. Read the triggering and acquisition periods of the synchronization echo sequence at the time level to make the candidate sector set correspond one-to-one with the corresponding time period; Based on the above three types of correspondence, a sampling strategy is configured for each candidate sector: when the overlap area between the candidate sector and the candidate seepage mask is large, dense sampling with a small step size is set; when the overlap area is small or located in the risk observation area, sparse sampling with a large step size is set, and the sampling reason is recorded simultaneously.

4. The method according to claim 1, characterized in that, The process of performing cross-modal alignment and consistency calculation with the initial sequence of structural detection values ​​also includes: After the sampling strategy is determined, the original records of the synchronous echo sequence are read along the distance and time directions. The changes in echo intensity, echo arrival position, interlayer continuity and reflection pattern are calculated. The results of multiple samplings within the same sector are denoised and merged to generate feature segments with time markers, sector markers and distance markers. The feature fragment is bound to the spatial index of the candidate water seepage mask, so that each radar feature can be traced back to the candidate region and time segment from which it originated. For regions that are removed or sparsely processed, the mask state and sampling step size are registered in the feature fragment; All feature segments are organized according to sector order and time order to obtain local radar feature sequences and output them.

5. The method according to claim 1, characterized in that, The process of performing cross-modal alignment and consistency calculation also includes: Establish a one-to-one correspondence between each feature segment in the local radar feature sequence and the candidate region bound to it; at the time level, based on the previously registered time stamps, compare segments from the radar side and the image side within the same candidate region in the same time segment. The apparent quantities related to boundary stability, texture change and local darkening in the initial value sequence of structural detection quantities are read and mapped to the local window of each candidate region. They are then compared item by item with the echo intensity change, arrival position change and interlayer continuity of the local radar feature sequence.

6. The method according to claim 1, characterized in that, The process of performing cross-modal alignment and consistency calculation also includes: For each comparison, a source identifier, region identifier, and time period identifier are attached to the recording end. Based on the comparison results, a hierarchical registration of comparison entries is generated within the spatial grid of the candidate region. First, the direction and magnitude of change on the image side and the radar side are matched within the frame, and then the continuity and repetition are matched between frames. Entries marked as risk observation areas by the region of interest are weighted down during registration. Entries registered as low confidence by the candidate water seepage mask are given source hints during registration. Finally, the above entries are summarized in each candidate region according to spatial grid and time segment to form a scoring record with weights and identifiers. The record is organized into a consistency scoring sequence according to the regional order and time order and output.

7. The method according to claim 1, characterized in that, The process of weighting and energy construction of voxels using the region of interest map as the outer frame discrete voxel set also includes: At the spatial level, using the region of interest map as the outer frame, the spatial range covered by the candidate region is discretized into a set of voxels conforming to the registration specifications of this system, and an index field from the image side and the radar side is established for each voxel. At the temporal level, using the time stamp registered in the previous order as a reference, the candidate state and consistency of the same voxel in adjacent frames are recorded as time entries. Then, weighting is carried out at the voxel level: image-side weighting is based on the candidate water penetration mask, and the candidate state of the voxel in the corresponding frame is registered as the image-side weight base value, and the base value is attenuated or maintained by the uncertainty map; physical-side weighting is based on the consistency score sequence, and the score records corresponding to the voxel in the same region are aggregated by time entries and registered as the physical-side weight base value.

8. The method according to claim 7, characterized in that, Also includes: When a voxel is located in a shieldable region or a risk observation area of ​​the region of interest map, the weight base values ​​of the image side and the physical side are uniformly weakened or delayed in effect according to the registration rules of that region, and the reasons for weakening or delaying are recorded.

9. The method according to claim 8, characterized in that, The process of generating a voxel energy map also includes: After completing the image-side and physical-side weighting, the weights of voxels on the time entries are checked for continuity. Time periods in which the same voxel is continuously candidate across frames and is continuously supported by the physical side are marked as continuous entries, while voxels that appear only in individual frames are marked as isolated entries. Different time penalty rules are registered for continuous entries and isolated entries at the recording end. After completing the above weighting, a unified energy registration value is formed for each voxel based on its image-side weight base value, physical-side weight base value, and time entry penalty record. The voxel set is then organized into a continuous voxel energy map according to the region and time order and output.

10. A bridge foundation seepage intelligent detection system, applied to the method according to any one of claims 1-9, characterized in that, include: The multimodal data acquisition and calibration module is used to acquire the raw response and calibration reference data of infrared cameras, visible light cameras and ground penetration radar, and to perform geometric calibration, radiometric calibration and unified time baseline establishment to generate joint calibration parameters. The synchronous acquisition and control module is used to calculate the synchronization relationship between excitation and sampling based on joint calibration parameters to generate adaptive thermal excitation parameters and phase-locked sampling parameters, and to generate synchronous image sequences and synchronous echo sequences based on the region of interest map for trigger configuration and coherent acquisition. The data alignment and preprocessing module is used to perform temporal and spatial alignment based on the synchronized image sequence and joint calibration parameters to generate an aligned image sequence, and to perform local cropping and pose verification based on the aligned image sequence to generate an initial value sequence of structural detection quantities. The drift estimation and compensation module is used to estimate temperature drift and illumination drift based on the initial value sequence of structural detection quantities to generate environmental drift parameters, perform radiation compensation and illumination compensation based on the environmental drift parameters to generate a compensated image sequence, and perform drift region labeling and noise modeling based on the compensated image sequence to generate a drift mask. The feature extraction and model module is used to extract and locally filter thermal and wet fingerprint features based on the compensated image sequence and drift mask. Based on the thermal and wet fingerprint features, the module trains the segmentation model to generate a pre-trained segmentation model. Based on the pre-trained segmentation model, the module performs inference to generate candidate seepage masks and uncertainty maps. The consistency verification module is used to generate local radar feature sequences by sector sampling and step size configuration based on candidate seepage masks, to generate consistency score sequences by consistency calculation based on local radar feature sequences, and to generate voxel energy maps by voxel weighting and energy construction based on consistency score sequences. The optimization and localization module is used to perform optimization and boundary refinement based on the voxel energy map to generate a refined mask sequence, and to perform spatial fitting and voxel clustering based on the refined mask sequence to generate a set of seepage areas and a spatial location sequence. The parameter management module is used to perform statistical analysis and parameter updates based on the set of seepage areas and spatial location sequences, generating parameter write-back packages.