An air-ground cooperative multi-modal industrial building durability detection system and method
By using a multimodal detection system that combines air and ground operations with drones and wall-climbing detection modules for 3D reconstruction and multidimensional data acquisition, the problem of false alarms and missed detections in heavy industrial plant inspections has been solved, achieving high-precision durability assessment and reliability testing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
- Filing Date
- 2026-04-14
- Publication Date
- 2026-07-10
AI Technical Summary
In the testing of heavy industrial plants with high temperature, high humidity, strong corrosion, and high dust, existing technologies are prone to interference with single-modal testing methods, leading to false alarms and missed detections. Furthermore, multi-source heterogeneous data cannot be synergistically integrated, making it difficult to achieve accurate durability assessment.
A multimodal detection system combining air and ground operations is adopted, which integrates UAVs and wall-climbing detection modules to perform three-dimensional reconstruction, multi-dimensional data acquisition and analysis. The multimodal analysis module enables data fusion and comprehensive judgment to generate a durability assessment report.
It achieves high-precision, reliable, and durable assessment under complex working conditions, adapts to the safety management requirements of heavy industrial plants, and improves the accuracy and reliability of testing.
Smart Images

Figure CN122367967A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of building structure testing technology, and relates to a multimodal industrial building durability testing system and method that combines air and ground operation. Background Technology
[0002] In heavy industries such as metallurgy, power, and chemicals, parts of factory buildings are subjected to prolonged exposure to a complex environment of high temperature, high humidity, and highly corrosive gases and dust. These conditions accelerate concrete carbonation, induce stress corrosion cracking in reinforcing steel, and lead to the destruction of the passivation film on the steel structure surface, the expansion of pitting, and large-scale blistering and peeling of the coating, severely weakening the structure's load-bearing capacity and service life.
[0003] Currently, in order to ensure the safe and stable operation of heavy industrial plant structures and avoid potential safety hazards, the industry generally uses drones equipped with visible light cameras or infrared thermal imagers to conduct high-altitude area inspections. This is combined with manual visual inspections, rebound hammer sampling tests, and simple chemical index tests to identify damage to plant structures, conduct preliminary assessments of physical and mechanical strength, and make simple determinations of the degree of chemical corrosion. Among these, physical and mechanical parameters are mainly obtained through equipment such as rebound hammers, while chemical corrosion indicators are used for simple detection of corrosion products on the structural surface and environmental corrosive media.
[0004] However, existing detection methods face significant technical bottlenecks in dealing with the harsh combined conditions of high temperature, high humidity, strong corrosion, and high dust levels. Firstly, single-modal detection methods have weak anti-interference capabilities and are easily obstructed by debris such as fly ash, slag, and corrosive stains accumulated on the surface of factory structures. This leads to frequent false alarms such as false cracks and false hot spots during detection, and also easily causes missed detections of actual damage, severely affecting the accuracy and reliability of the test results. Secondly, multi-source heterogeneous detection data, such as visible light images, infrared thermography, physical and mechanical parameters, and chemical corrosion indicators, have long been fragmented. Each modal data is independent and cannot form a synergistic correlation, making it difficult to combine with knowledge from fields such as industrial corrosion mechanisms and structural damage evolution laws for comprehensive reasoning. Consequently, it is impossible to achieve accurate judgment and quantitative assessment of the durability of factory structures, failing to meet the high-precision and high-reliability requirements of heavy industry for building structure safety management, thus hindering the upgrading and development of industrial building durability assessment technology. Summary of the Invention
[0005] The purpose of this invention is to provide a multimodal industrial building durability testing system and method that combines air and ground operations, enabling efficient acquisition, deep fusion, and accurate analysis of multimodal testing data under air-ground collaborative operation, thereby achieving rapid, accurate assessment and reliable control of industrial building durability.
[0006] To achieve the above objectives, the technical solution provided by the present invention is as follows: A multimodal industrial building durability testing system with air-ground collaboration includes: The air-ground collaborative perception module is used to perform 3D reconstruction of the building interior and simultaneously collect image data of the building structure surface. The 3D reconstruction is used to obtain a 3D environment map of the building interior. The image data of the building structure surface is used to identify defects and obtain the planar position of the suspected defect area of the building structure. Combining the 3D environment map and the real-time pose information when the images were collected, the planar position of the suspected defect area is mapped to the 3D environment map coordinate system to determine the 3D spatial coordinates of the suspected defect area and generate a detection task list containing the 3D coordinates of the suspected defect area. The wall-climbing detection module is used to receive the detection task list, autonomously navigate to the target area based on the three-dimensional spatial coordinates of the suspected defect area in the detection task list, and perform close-range operations. It operates in an orderly manner according to the progressive process of surface cleaning, multi-dimensional physical detection, and micro-destructive physical sampling, and simultaneously collects visible light images, infrared thermograms, physical hardness parameters, geometric parameters and in-situ ion concentration titration data of the target area. The multimodal analysis module is used to receive 3D environmental maps, building structure surface image data, planar location and 3D spatial coordinates of suspected defect areas, visible light images, infrared thermograms, physical hardness parameters, geometric parameters and in-situ ion concentration titration data, and process and analyze them to obtain an industrial building structure durability assessment report that includes damage type, degree, distribution and durability level.
[0007] The invention is further characterized by: The air-ground collaborative perception module includes: a drone for autonomous flight inside the building and acquisition of real-time pose information during its flight; a lidar unit mounted on the drone for collecting environmental point cloud data inside the building; an industrial camera mounted on the drone for synchronously acquiring image data of the building's structural surface; and a processing unit located inside the drone for receiving environmental point cloud data inside the building and image data of the building's structural surface. This processing unit processes the environmental point cloud data inside the building to generate laser SLAM front-end odometry information, uses this information to correct the drone's real-time pose information, completes the 3D reconstruction of the building's interior, generates a globally consistent 3D environmental map, and simultaneously identifies defects in the image data of the building's structural surface, obtaining the planar positions of suspected defect areas. These planar positions are then mapped to the 3D environmental map coordinate system to determine the 3D spatial coordinates of the suspected defect areas, generating a detection task list containing the 3D coordinates of the suspected defect areas.
[0008] When the processing unit processes the environmental point cloud data inside the building, it sequentially performs timestamp synchronization, motion distortion removal, point cloud registration, and cumulative error elimination on the environmental point cloud data inside the building, and calculates the relative pose between frames in real time during the point cloud registration process.
[0009] The air-ground collaborative perception module also includes a visualization terminal, which receives a 3D environment map. This allows engineers to mark the 3D coordinates of specific areas of interest on the 3D environment map based on historical maintenance records, original design drawings, and on-site visual judgment. The map is then sent to the processing unit, which merges the 3D coordinates of suspected defect areas with the 3D coordinates of specific areas of interest. Simultaneously, it prioritizes the areas based on their importance and defect confidence, generating a list of detection tasks that includes the priority ranking.
[0010] The air-ground cooperative perception module also includes a hierarchical navigation planning unit, which is located inside the UAV. The hierarchical navigation planning unit is used to fuse the laser SLAM front-end odometry information with the real-time pose information of the UAV after correction using the laser SLAM front-end odometry information to obtain the real-time pose information of the image data of the building structure surface. Combined with the 3D environment map, it generates the optimal navigation trajectory for the UAV that meets the kinematic constraints.
[0011] The wall-climbing detection module includes: a negative pressure adsorption walking unit for walking on the surface of the building structure; a cleaning unit for cleaning the target area; a dual-spectrum visual acquisition unit for acquiring visible light images and infrared thermal images of the target area; a digital rebound unit for acquiring physical hardness parameters of the target area; a line laser scanning unit for acquiring geometric parameters of the target area; an ion concentration detection unit for acquiring in-situ ion concentration titration data of the target area; a micro-destructive physical sampling unit for performing micro-destructive physical sampling of the target area; and a control unit for receiving the detection task list, controlling the negative pressure adsorption walking unit to autonomously navigate to the target area to perform close-range operations based on the three-dimensional spatial coordinates of suspected defect areas in the detection task list, and controlling the cleaning unit, dual-spectrum visual acquisition unit, digital rebound unit, line laser scanning unit, ion concentration detection unit, and micro-destructive physical sampling unit to operate in an orderly manner according to the progressive process.
[0012] The multimodal analysis module includes a data processing unit that receives 3D environmental maps, building structure surface image data, the planar location and 3D spatial coordinates of suspected defect areas, visible light images, infrared thermograms, physical hardness parameters, geometric parameters, and in-situ ion concentration titration data. It extracts texture features from the building structure surface image data and visible light images, temperature field features from the infrared thermograms, geometric topological features from the geometric parameters, and numerical sequence features from the physical hardness parameters and in-situ ion concentration titration data. Then, through a cross-attention mechanism, it maps the features of different modalities to a unified high-dimensional semantic space for alignment and fusion. Combining the 3D environmental map and the planar location and 3D spatial coordinates of suspected defect areas, it constructs a multimodal embedding vector that comprehensively represents the health status of the building structure. Analysis The decision-making unit is used to pre-build a knowledge base containing building structure durability design specifications, industrial corrosion mechanism literature, and historical testing cases. A large language model is trained using these knowledge bases. Multimodal embedding vectors are input into the large language model, which then searches and retrieves the most relevant standard provisions or similar historical testing cases from the knowledge base in real time. These are used as contextual prompts to provide the large language model with standard bases and case references for durability assessment. This assists the large language model in conducting comprehensive reasoning analysis, completing the analysis and diagnosis of the durability status of industrial buildings, and generating a durability assessment report for industrial building structures that includes damage type, degree, distribution, and durability level.
[0013] When the data processing unit extracts the texture features of the building structure surface image data and visible light image, the temperature field features of the infrared thermogram, the geometric topological features of the geometric parameters, and the numerical sequence features of the physical hardness parameters and in-situ ion concentration titration data, it uses a visual converter to extract the texture features of the building structure surface image data and visible light image, uses a point cloud network++ to extract the geometric topological features of the geometric parameters, and uses a multilayer perceptron to extract the temperature field features of the infrared thermogram and the numerical sequence features of the physical hardness parameters and in-situ ion concentration titration data.
[0014] When the data processing unit maps features from different modalities to a unified high-dimensional semantic space for alignment and fusion through a cross-attention mechanism, it employs a multi-source heterogeneous feature alignment technique.
[0015] A multimodal durability testing method for air-ground collaborative industrial buildings includes the following steps: The interior of the building is reconstructed in 3D and the surface image data of the building structure are collected simultaneously. The 3D reconstruction is used to obtain a 3D environment map of the interior of the building. The surface image data of the building structure is used to identify defects and obtain the planar position of the suspected defect area of the building structure. Combining the 3D environment map and the real-time pose information when the images are collected, the planar position of the suspected defect area is mapped to the coordinate system of the 3D environment map to determine the 3D spatial coordinates of the suspected defect area and generate a detection task list containing the 3D coordinates of the suspected defect area. Based on the three-dimensional spatial coordinates of the suspected defective areas in the inspection task list, the system autonomously navigates to the target area to perform close-range operations. It operates in an orderly manner according to the progressive process of surface cleaning, multi-dimensional physical detection, and micro-destructive physical sampling, while simultaneously collecting visible light images, infrared thermal images, physical hardness parameters, geometric parameters, and in-situ ion concentration titration data of the target area. The process involves processing and analyzing 3D environmental maps, building structure surface image data, planar locations and 3D spatial coordinates of suspected defect areas, visible light images, infrared thermograms, physical hardness parameters, geometric parameters, and in-situ ion concentration titration data to obtain an industrial building structure durability assessment report that includes damage type, degree, distribution, and durability level.
[0016] The air-ground collaborative multimodal industrial building durability testing system and method of the present invention has the following advantages: This invention accurately reconstructs and conducts full-area inspection of the internal three-dimensional environment of a building, precisely locates suspected defect areas, and generates a list of inspection tasks containing the three-dimensional coordinates of these areas. Through a progressive, close-range operation using a wall-climbing inspection module, it completes multi-dimensional, multi-scenario, multi-source inspection data collection. Relying on a multi-modal analysis module, it achieves collaborative fusion and comprehensive analysis of various heterogeneous inspection data, accurately generating a durability assessment report for industrial building structures that includes damage type, degree, distribution, and durability level. This invention is adaptable to the harsh and complex working conditions of heavy industrial plants with high temperature, high humidity, strong corrosion, and high dust, providing high-precision and high-reliability durability testing and assessment support for the safety management of industrial building structures, and contributing to the upgrading and development of industrial building durability assessment technology. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the overall structure of the present invention.
[0018] Figure 2 This is a schematic diagram of the overall process of the present invention.
[0019] Figure 3 This is a flowchart of the wall-climbing detection module in this invention.
[0020] Figure 4 This is a flowchart of the dual-spectral vision acquisition unit in this invention.
[0021] Figure 5 This is a flowchart of the multimodal analysis module in this invention. Detailed Implementation
[0022] The technical solutions of the present invention will now be described clearly and in detail with reference to the accompanying drawings. In the description of the embodiments of the present invention, unless otherwise stated, " / " indicates "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, in the description of the embodiments of the present invention, "multiple" refers to two or more. The terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature.
[0023] like Figure 1As shown, this invention provides an air-ground collaborative multimodal industrial building durability testing system, including an air-ground collaborative sensing module, a wall-climbing detection module, and a multimodal analysis module. The air-ground collaborative sensing module is used to perform 3D reconstruction of the building interior and simultaneously acquire image data of the building structure surface. The 3D reconstruction generates a 3D environment map of the building interior, and the building structure surface image data is used for defect identification to obtain the planar location of suspected defect areas. Combining the 3D environment map and real-time pose information during image acquisition, the planar location of the suspected defect areas is mapped to the 3D environment map coordinate system to determine the 3D spatial coordinates of the suspected defect areas, generating a detection task list containing the 3D coordinates of the suspected defect areas. The wall-climbing detection module is used to receive... The system autonomously navigates to the target area based on the 3D spatial coordinates of suspected defective areas in the inspection task list to perform close-range operations. It operates in an orderly manner according to the progressive process of surface cleaning, multi-dimensional physical detection, and micro-destructive physical sampling. Simultaneously, it collects visible light images, infrared thermograms, physical hardness parameters, geometric parameters, and in-situ ion concentration titration data of the target area. The multimodal analysis module receives 3D environmental maps, building structure surface image data, the planar location and 3D spatial coordinates of suspected defective areas, visible light images, infrared thermograms, physical hardness parameters, geometric parameters, and in-situ ion concentration titration data, processes and analyzes them, and obtains an industrial building structure durability assessment report that includes damage type, degree, distribution, and durability level. This invention accurately reconstructs and conducts full-area inspection of the internal three-dimensional environment of a building, precisely locates suspected defect areas, and generates a list of inspection tasks containing the three-dimensional coordinates of these areas. Through a progressive, close-range operation using a wall-climbing inspection module, it completes multi-dimensional, multi-scenario, multi-source inspection data collection. Relying on a multi-modal analysis module, it achieves collaborative fusion and comprehensive analysis of various heterogeneous inspection data, accurately generating a durability assessment report for industrial building structures that includes damage type, degree, distribution, and durability level. This invention is adaptable to the harsh and complex working conditions of heavy industrial plants with high temperature, high humidity, strong corrosion, and high dust, providing high-precision and high-reliability durability testing and assessment support for the safety management of industrial building structures, and contributing to the upgrading and development of industrial building durability assessment technology.
[0024] like Figure 1As shown, the air-ground collaborative perception module includes a drone, a LiDAR, an industrial camera, and a processing unit. The drone is used to autonomously fly inside the building and acquire its real-time pose information during flight. The LiDAR is mounted on the drone and is used to collect environmental point cloud data inside the building. The industrial camera is mounted on the drone and is used to synchronously acquire image data of the building's structural surface. The processing unit is located inside the drone and is used to receive the environmental point cloud data inside the building and the image data of the building's structural surface. After processing the environmental point cloud data inside the building, it generates LiDAR SLAM front-end odometry information. The LiDAR SLAM front-end odometry information is used to correct the drone's real-time pose information, complete the 3D reconstruction of the building's interior, and generate a globally consistent 3D environmental map. At the same time, it performs defect identification on the image data of the building's structural surface, obtains the planar position of suspected defect areas of the building structure, maps the planar position of the suspected defect areas to the 3D environmental map coordinate system, determines the 3D spatial coordinates of the suspected defect areas, and generates a detection task list containing the 3D coordinates of the suspected defect areas.
[0025] like Figure 1 As shown, when the processing unit processes the environmental point cloud data inside the building, it sequentially performs timestamp synchronization, motion distortion removal, point cloud registration, and cumulative error elimination on the environmental point cloud data inside the building, and calculates the inter-frame relative pose in real time during the point cloud registration process.
[0026] Specifically, the processing unit performs timestamp synchronization and motion distortion removal on the environmental point cloud data acquired by the lidar, calculates the relative pose between frames using a point cloud registration algorithm, outputs high-update-frequency front-end odometry information, selects key frames, performs back-end optimization processing based on the distorted point cloud, eliminates accumulated errors by constructing a factor graph or graph optimization model, constructs a globally consistent environmental point cloud map, and outputs low-update-frequency high-precision pose information. The front-end odometry information is used to ensure the real-time performance of UAV attitude control, and the back-end optimization is used to ensure the accuracy of global positioning.
[0027] like Figure 1 As shown, the air-ground collaborative sensing module also includes a visualization terminal. The visualization terminal is used to receive a 3D environment map, which allows engineers to mark the 3D coordinates of specific areas of interest in the 3D environment map based on historical maintenance records, original design drawings, and on-site intuitive judgment. The map is then sent to the processing unit, which merges the 3D coordinates of suspected defect areas with the 3D coordinates of specific areas of interest. At the same time, it prioritizes the areas based on their importance and defect confidence, and generates a list of detection tasks that includes the priority ranking.
[0028] like Figure 1As shown, the air-ground cooperative perception module also includes a hierarchical navigation planning unit, which is located inside the UAV. The hierarchical navigation planning unit is used to fuse the laser SLAM front-end odometry information with the real-time pose information of the UAV after correction using the laser SLAM front-end odometry information to obtain the real-time pose information of the image data of the building structure surface. Combined with the 3D environment map, it generates the optimal navigation trajectory for the UAV that meets the kinematic constraints.
[0029] Among them, SLAM stands for Simultaneous Localization and Mapping Front End.
[0030] When the hierarchical navigation planning unit fuses the laser SLAM front-end odometry information with the UAV's real-time pose information corrected using the laser SLAM front-end odometry information, it uses extended Kalman filtering or optimization algorithms.
[0031] In this process, when the hierarchical navigation planning unit generates the optimal navigation trajectory for the UAV that satisfies the kinematic constraints by combining the three-dimensional environment map, it uses a hybrid A* algorithm to search and generate an initial collision-free path that satisfies the UAV's kinematic constraints. The key nodes of this initial path are used as control points, and the trajectory is parametrically fitted using B-spline curves to construct a trajectory optimization function that includes smoothness cost, obstacle distance cost, and dynamic feasibility cost. This generates the optimal navigation trajectory, thereby enabling the UAV to achieve stable hovering and autonomous obstacle avoidance flight inside complex buildings with no GPS signal and dust interference.
[0032] When the processing unit performs defect identification on the image data of the building structure surface, it uses a target detection algorithm to perform real-time reasoning analysis on the image data of the building structure surface, automatically identifying and locating appearance defects such as cracks, peeling, exposed reinforcement and leakage. The target detection algorithm is preferably the YOLO algorithm.
[0033] To further ensure the comprehensiveness and specificity of the inspection and avoid safety hazards caused by the processing unit's omission of defects in image data, the processing unit also has a pre-set mandatory inspection mechanism based on the structural importance level. The structural importance level refers to the priority classification of various components, taking into account their functional positioning in the overall structural force transmission system, the impact of failure on structural safety and usability, the severity of their service environment, and the availability of alternatives for subsequent inspections. Based on the general principles of structural inspection and safety assessment of old industrial buildings, key components in the main load-bearing frame system, such as main beams, columns, beam-column joints, roof truss joints, supporting components, column bases, and foundations, should generally be classified as high importance because they directly bear the main load transmission function and have a significant impact on the overall force path and structural unit safety. For components that have been in adverse service environments such as high temperature, corrosion, humidity, vibration, or impact for a long time, even if they are not currently identified by the processing unit through image data collected by industrial cameras as having obvious appearance defects, they may pose potential risks due to material performance degradation or load-bearing capacity reduction; therefore, their structural importance level should also be increased accordingly. This grading approach is consistent with the requirements in the "Industrial Building Reliability Appraisal Standard" (GB50144-2019) that industrial buildings should be appraised based on the importance of their components, and the technical approach of implementing safety assessments for old industrial buildings in a step-by-step manner according to "components - structural units - structural systems".
[0034] Based on the aforementioned mandatory inspection mechanism, when the processing unit identifies main beam or column joints bearing significant loads, or critical component areas operating under adverse conditions such as high temperatures or corrosion, the system automatically records the three-dimensional center coordinates of the corresponding component and marks it as a mandatory inspection point, regardless of whether the image data collected by the processing unit through the industrial camera identifies any external defects. This is then incorporated into the subsequent refined inspection process, working in conjunction with the suspected defect areas and specific areas of interest marked by engineers mentioned earlier to improve the comprehensiveness of the inspection task list. The aforementioned mandatory inspection points refer to spatial positioning points that require priority verification and supplementary inspection during the inspection task, ensuring that critical load-bearing components are not missed due to a single processing unit's normal image data recognition result. The aforementioned "three-dimensional center coordinates" refer to the component's positional identifier in the building space model or inspection coordinate system, serving as a reference information for subsequent automatic navigation of inspection equipment, inspection data binding, and time-series comparison of defects. For these mandatory inspection points, subsequent inspection items can be automatically selected based on the component type and according to specifications; for example, for concrete components, further inspections can be conducted on strength, cracks, carbonation depth, rebar location and corrosion, deflection, or displacement.
[0035] like Figure 1As shown, the wall-climbing detection module includes a negative pressure adsorption walking unit, a cleaning unit, a dual-spectrum visual acquisition unit, a digital rebound unit, a line laser scanning unit, an ion concentration detection unit, a micro-destructive physical sampling unit, and a control unit. The negative pressure adsorption walking unit is used to walk on the surface of the building structure; the cleaning unit is used to clean the target area; the dual-spectrum visual acquisition unit is used to acquire visible light images and infrared thermal images of the target area; the digital rebound unit is used to acquire physical hardness parameters of the target area; the line laser scanning unit is used to acquire geometric parameters of the target area; the ion concentration detection unit is used to acquire in-situ ion concentration titration data of the target area; the micro-destructive physical sampling unit is used to perform micro-destructive physical sampling of the target area; and the control unit is used to receive the detection task list, control the negative pressure adsorption walking unit to autonomously navigate to the target area to perform close-range operations based on the three-dimensional spatial coordinates of the suspected defect area in the detection task list, and control the cleaning unit, dual-spectrum visual acquisition unit, digital rebound unit, line laser scanning unit, ion concentration detection unit, and micro-destructive physical sampling unit to operate in an orderly manner according to the progressive process.
[0036] The wall-climbing detection module also includes a biomimetic microstructure adsorption unit and a pneumatic stabilization unit. The control unit is electrically connected to the biomimetic microstructure adsorption unit and the pneumatic stabilization unit respectively. On the one hand, when the wall-climbing detection module reaches the target area and performs fixed-point operation, the control unit first controls the biomimetic microstructure adsorption unit to adhere tightly to the structural surface, using van der Waals force or micro-interlocking effect to provide basic tangential anti-slip resistance. On the other hand, strictly following the graded detection process, when rebound detection or micro-destructive physical sampling tasks need to be performed, the pneumatic stabilization unit is pre-controlled to start before the contact action occurs, outputting high-pressure airflow in a directional direction away from the wall surface, using pneumatic thrust to apply active normal clamping force to the wall-climbing detection module, achieving strong adhesion and attitude locking between it and the building structure surface. Through the synergistic effect of the above-mentioned "static anchoring of the biomimetic microstructure adsorption unit" and "active clamping of the task-triggered pneumatic stabilization unit", a temporary quasi-static working base is constructed on the vertical or inverted surface, ensuring the stability of high-precision sensing and sampling operations.
[0037] The control unit incorporates cascaded condition-triggered detection logic based on damage feature confidence levels. After each detection step, it continuously assesses the abnormal characteristics of the current data and performs logical judgments to dynamically determine whether to trigger the next level of detection. The specific process is as follows: First, a non-contact scanning is performed using a dual-spectrum visual acquisition unit. Only when the feature confidence level of a suspected damaged area exceeds a preset threshold is a command generated to trigger the wall-climbing detection module to activate the cleaning component and remove surface dust. After the cleaning operation is completed, the dual-spectrum visual acquisition unit immediately performs visual verification. If the original suspected area is determined to be a false positive, the task at the current location is automatically terminated. If it is confirmed to be actual structural damage, the digital rebound unit and line laser scanning unit are further triggered to perform intensity detection and precision geometric scanning. Furthermore, only when the above multimodal data further points to potential internal chemical corrosion risks is the micro-damage physical sampling unit triggered to carry out operations. This process realizes a progressive and on-demand operation from "non-contact initial screening" to "contact verification" and then to "micro-damage diagnosis," significantly improving operational efficiency while ensuring detection accuracy.
[0038] The dual-spectral visual acquisition unit is based on a physically consistent prior dual-spectral enhanced detection method. It includes a visible light imaging device and an infrared thermal imaging device. The control unit first performs spatiotemporal alignment and radiometric correction processing on the acquired dual-spectral data. Specifically, it uses a pre-calibrated homography matrix to map the infrared thermal image to the visible light image coordinate system, achieving sub-pixel-level spatial registration. Simultaneously, to address temperature measurement deviations caused by industrial dust, the system introduces a dark channel prior algorithm to analyze the visible light image to calculate the wall dust concentration distribution map. Based on this, it dynamically compensates the emissivity parameters of the infrared thermal image pixel by pixel, reconstructing a true surface temperature field image that eliminates surface dust thermal resistance interference and reflects the actual heat flow characteristics inside the structure. Before deep learning inference, the control unit introduces a physical logic gating mechanism to generate a "physical prior attention mask" to guide subsequent feature extraction. Specifically, the Sobel edge detection operator is used to calculate the texture gradient of the visible light image, while simultaneously calculating the temperature gradient of the real surface temperature field image. Physical consistency logic is then applied: for regions with strong visible light texture gradients but no significant temperature gradient, the system classifies them as "surface stains" or "pseudo-cracks" and generates a suppression mask; for regions with weak visible light textures but significant temperature gradients, the system classifies them as "hidden voids" or "deep cracks" and generates an enhancement mask. This mechanism encodes the physical laws of heat transfer into logical constraints, effectively filtering visual artifact noise. The control unit also constructs a dual-stream interactive neural network based on an improved CBAM (Convolutional Block Attention Module) for joint inference on the corrected dual-spectral data. The network includes parallel RGB stream encoders and thermal feature encoders. Specifically, the feature map extracted by the thermal encoder is pooled and then input into the dual-spectral image gating unit. This unit combines the physical prior attention mask to generate a thermal saliency weight matrix, which is then injected into the RGB stream encoder as a gating signal. By updating the weights, the visible light feature map is weighted and modulated, forcing the network to focus on structural damage areas with real thermal anomalies. While the dual-spectral data fusion decoding outputs semantic segmentation results, the control unit performs damage depth quantization estimation in parallel. Specifically, the control unit extracts the maximum temperature gradient features within the damaged area based on the real surface temperature field image, and combines it with a pre-set concrete heat conduction inversion model to calculate the depth of voids or cracks. Finally, it outputs a comprehensive detection result containing damage type, planar geometric parameters, and depth information. When the visible light image analysis determines that the damage severity exceeds a preset threshold, the dual-spectral vision processing unit automatically triggers the digital rebound unit to perform strength verification. Since rebound detection is a high-contact impact operation, the system first activates the pneumatic stabilization unit before the impact action, using reverse airflow to press the robot body against the working surface.Subsequently, the rebound detection unit automatically aligns the probe with the working surface via an electrically driven attitude adjustment mechanism. It utilizes a built-in high-precision photoelectric or magnetic grating displacement sensor to collect real-time impact and rebound displacement data of the impact rod, converting the mechanical motion into digital signals and uploading them. It automatically records the rebound sequence at multiple points in the test area and combines this with the attitude tilt information from the wall-climbing detection module for energy correction, generating the estimated concrete strength data for that point. If the rebound strength is lower than the preset structural safety threshold, a command is generated to trigger the next level of high-precision geometric scanning.
[0039] When the rebound strength of the rebound detection unit shows abnormal feedback, the control unit activates the linear laser scanning unit to perform non-contact micro-geometric reconstruction of the target damaged area. The control unit uses the principle of triangulation to obtain a high-resolution three-dimensional point cloud of the crack or spalling pit, and calculates the fine geometric parameters of the damage (including the average crack width, maximum depth, opening area, and surface roughness) through point cloud topology analysis algorithms. These geometric quantification indicators are not only used to assess the current degree of physical damage, but also serve as a key decision-making basis. If the scanning results show that the crack depth or opening area indicates a risk of deep corrosion, the control unit will ultimately trigger the micro-damage physical sampling component to enter the deepest level of micro-damage physical sampling stage.
[0040] The micro-destructive physical sampling unit is activated after confirming the need for chemical analysis. The control unit then confirms that the pneumatic stabilization unit is in a high-power state to resist grinding reaction forces. Subsequently, samples are taken from the target area to obtain deep concrete powder samples.
[0041] like Figure 1As shown, the multimodal analysis module includes a data processing unit and an analysis and decision-making unit. The data processing unit receives 3D environmental maps, building structure surface image data, the planar location and 3D spatial coordinates of suspected defect areas, visible light images, infrared thermograms, physical hardness parameters, geometric parameters, and in-situ ion concentration titration data. It extracts texture features from the building structure surface image data and visible light images, temperature field features from the infrared thermograms, geometric topological features from the geometric parameters, and numerical sequence features from the physical hardness parameters and in-situ ion concentration titration data. Then, through a cross-attention mechanism, the features of different modalities are mapped to a unified high-dimensional semantic space for alignment and fusion. Combining the 3D environmental map and the planar location and 3D spatial coordinates of suspected defect areas, a multimodal analysis is constructed to comprehensively characterize the health status of the building structure. The embedding vector analysis and decision-making unit is used to pre-build a knowledge base containing building structure durability design specifications, industrial corrosion mechanism literature, and historical testing cases. A large language model is trained using these knowledge bases. Multimodal embedding vectors are input into the large language model, which then searches and retrieves the most relevant standard provisions or similar historical testing cases from the knowledge base in real time. These are used as contextual prompts to provide the large language model with standard bases and case references for durability assessment. This assists the large language model in conducting comprehensive reasoning analysis, completing the analysis and diagnosis of the durability status of industrial buildings, and generating a durability assessment report for industrial building structures that includes damage type, degree, distribution, and durability level.
[0042] like Figure 1 As shown, when the data processing unit extracts the texture features of the building structure surface image data and visible light image, the temperature field features of the infrared thermogram, the geometric topological features of the geometric parameters, and the numerical sequence features of the physical hardness parameters and in-situ ion concentration titration data, it uses a visual transformer to extract the texture features of the building structure surface image data and visible light image, uses a point cloud network++ (PointNet++) to extract the geometric topological features of the geometric parameters, and uses a multilayer perceptron (MLP) to extract the temperature field features of the infrared thermogram and the numerical sequence features of the physical hardness parameters and in-situ ion concentration titration data.
[0043] When the data processing unit maps features from different modalities to a unified high-dimensional semantic space for alignment and fusion through a cross-attention mechanism, it employs a multi-source heterogeneous feature alignment technique.
[0044] This invention also provides a multimodal industrial building durability testing method with air-ground collaboration, comprising the following steps: The interior of the building is reconstructed in 3D and the surface image data of the building structure are collected simultaneously. The 3D reconstruction is used to obtain a 3D environment map of the interior of the building. The surface image data of the building structure is used to identify defects and obtain the planar position of the suspected defect area of the building structure. Combining the 3D environment map and the real-time pose information when the images are collected, the planar position of the suspected defect area is mapped to the coordinate system of the 3D environment map to determine the 3D spatial coordinates of the suspected defect area and generate a detection task list containing the 3D coordinates of the suspected defect area. Based on the three-dimensional spatial coordinates of the suspected defective areas in the inspection task list, the system autonomously navigates to the target area to perform close-range operations. It operates in an orderly manner according to the progressive process of surface cleaning, multi-dimensional physical detection, and micro-destructive physical sampling, while simultaneously collecting visible light images, infrared thermal images, physical hardness parameters, geometric parameters, and in-situ ion concentration titration data of the target area. The process involves processing and analyzing 3D environmental maps, building structure surface image data, planar locations and 3D spatial coordinates of suspected defect areas, visible light images, infrared thermograms, physical hardness parameters, geometric parameters, and in-situ ion concentration titration data to obtain an industrial building structure durability assessment report that includes damage type, degree, distribution, and durability level.
[0045] Example 1 like Figure 1 , Figure 2 , Figure 3 As shown in the figure, this embodiment describes in detail the operation process of the detection system proposed in this invention in a real industrial building scenario. The entire operation process is not a simple linear execution, but adopts a cascaded decision logic of "vision-bounce-scanning-sampling" to ensure that the detection accuracy is maximized while minimizing the operation energy consumption.
[0046] At the start of the operation, the UAV in the air-ground collaborative perception module was activated first, performing a rapid scanning task inside the building where there was no GPS signal. The UAV's onboard LiDAR collected environmental point cloud data at high frequency and rotation speed. Based on the existing tightly coupled LiDAR SLAM framework, the inter-frame pose and odometry information were obtained through the front-end odometry and input into the back-end factor map for joint optimization, reconstructing a 3D environmental map of the building's interior. This provides a unified global spatial reference for subsequent operations. During flight, the airborne industrial camera simultaneously collected image data of the building's structural surface. When the processing unit processed the environmental point cloud data inside the building, it sequentially performed time-stamp synchronization and motion distortion correction on the environmental point cloud data inside the building. The process involves removing, registering, and eliminating accumulated errors in the point cloud. During point cloud registration, inter-frame relative pose is calculated in real-time. The environmental point cloud data inside the industrial building is processed to generate laser SLAM front-end odometry information. This information is used to correct the UAV's real-time pose, completing the 3D reconstruction of the industrial building's interior and generating a globally consistent 3D environment map. Simultaneously, defect identification is performed on the image data of the building structure's surface, such as cracks, leaks, or spalling features on beams, slabs, and columns. The planar positions of suspected defective areas are obtained and mapped to the 3D environment map coordinate system to determine their 3D spatial coordinates. Meanwhile, the UAV's hierarchical navigation planning unit fuses the laser SLAM front-end odometry information with the corrected real-time pose information of the UAV to obtain the real-time pose information of the building structure's surface image data. Combined with the 3D environment map, this generates an optimal navigation trajectory for the UAV that satisfies kinematic constraints. Meanwhile, the system operates a "human-machine collaborative task definition" mechanism, allowing engineers to manually mark the coordinates of specific areas of interest directly on a pre-constructed 3D environmental map coordinate system using a visual terminal, based on historical maintenance records, original design drawings, or intuitive on-site judgment. Finally, the suspected damage points perceived autonomously by the UAV and the supplementary points manually marked by engineers are summarized and deduplicated to generate a refined inspection task list containing priority ranking and precise spatial location, which is then distributed to the wall-climbing detection module via wireless LAN.
[0047] In response to the pre-processing command, the wall-climbing detection module activates the cleaning unit to physically remove dust and loose materials from the target area, exposing the true surface of the concrete. After cleaning, a secondary visual verification is immediately performed. If damage is confirmed, the process enters the strong-contact verification stage. At this point, to resist the reaction force generated by subsequent operations, the pneumatic stabilization unit is pre-activated, outputting high-pressure airflow away from the wall. Utilizing the combined action of pneumatic thrust and biomimetic adsorption components, the wall-climbing detection module is firmly pressed and locked to the building structure surface, constructing a quasi-static rigid working base. With a stable posture, the digital rebound unit collects and uploads rebound displacement data in real time. Subsequently, the control unit calculates the estimated concrete strength in real time based on the collected rebound displacement data. Only when the calculation result is lower than the preset structural safety threshold or exhibits abnormal strength dispersion is a further geometric quantification command generated. The line laser scanning unit is activated to perform micron-level three-dimensional reconstruction of cracks or spalling pits, obtaining geometric parameters such as crack width, depth, and opening area.
[0048] As the final step in the inspection process, a comprehensive defect assessment is performed based on the aforementioned visible light images, infrared thermograms, rebound strength, and geometric parameters. Specifically, the system fuses the geometric morphology of the spalling pits acquired by line laser scanning with the surface image data of the building structure extracted by the dual-spectrum visual acquisition unit and the texture features (i.e., the color features of corrosion products on the building structure surface) of the visible light images. The system focuses on identifying typical chemical corrosion characteristics such as rust expansion and cracking caused by steel reinforcement corrosion, the exudation of reddish-brown corrosion products, and the layered peeling of the protective layer. If the multi-source data fusion assessment results show that the concrete strength is significantly lower than the standard threshold, or that the crack depth is highly correlated with the rust and spalling characteristics of the building surface, it indicates a risk of deep chloride ion corrosion or sulfate corrosion within the structure, automatically triggering a micro-destructive physical sampling task.
[0049] At this point, the wall-climbing detection module maintains the high-power operation of its aerodynamic stabilization unit and uses the micro-damage physical sampling unit to sample the core areas of the damage (such as the deepest rust or the boundary of spalling). During this process, the module's current global coordinates (X, Y, Z) are simultaneously read, generating a unique electronic tag containing location information, sampling depth, and a timestamp. After completing all the above steps, the wall-climbing detection module resumes its cruise mode and proceeds to the next target point on the task list until the entire plant inspection is completed. Finally, all collected multimodal data is aggregated to a cloud server for subsequent feature alignment and analysis by the multimodal analysis module, generating an industrial building structure durability assessment report that includes damage type, degree, distribution, and durability level.
[0050] Example 2 like Figure 4As shown in the figure, this embodiment elaborates on the dual-spectral refined detection algorithm process of the edge computing unit and cloud server running on the wall climbing detection module, aiming to solve the detection distortion problem caused by uneven lighting, dust accumulation and background thermal radiation in harsh industrial environments.
[0051] First, in the multispectral data preprocessing stage, spatiotemporal alignment and physical correction are performed. Given the physical parallax between visible light and infrared thermal imaging, a pre-calibrated homography matrix is used to map the acquired infrared thermal image to the visible light image coordinate system, achieving sub-pixel-level spatial registration. Simultaneously, addressing the issue of infrared thermometry distortion caused by fly ash adhering to industrial walls, the algorithm introduces the Dark Channel Prior principle. It analyzes the visible light image to calculate the fly ash concentration distribution map of the wall, and based on this, performs pixel-by-pixel dynamic compensation on the emissivity parameters of the infrared thermal image. This reconstructs a "true surface temperature field image" that eliminates the thermal resistance interference of fly ash and reflects the true heat flow characteristics inside the structure, thereby effectively suppressing false hot spots caused by uneven fly ash thickness.
[0052] Secondly, in the pre-feature extraction stage, the system constructs a physically consistent prior mask (Physics-PriorMask) to introduce logical constraints. Specifically, the system uses the Sobel operator to calculate the texture gradient of the visible light image, while simultaneously calculating the heat flux gradient of the real surface temperature field image, and performs pixel-level physical logic classification: for regions with only strong visible light texture but no obvious temperature gradient, they are classified as "surface stains" and a suppression mask is generated; for regions with weak visible light texture but significant temperature gradient, they are classified as "hidden voids" and an enhancement mask is generated. This mask encodes the physical laws of heat transfer into spatial attention weights, providing clear prior guidance for subsequent deep learning networks.
[0053] Furthermore, the system constructs a dual-stream interactive convolutional neural network (DMGN) with a cross-modal spatial gating unit (CMSGU) for joint inference. The network includes an RGB stream encoder with ResNet-50 as the backbone and a lightweight thermal stream encoder. To achieve the mechanism of "guiding visual attention with thermal features," a CMSGU based on an improved CBAM architecture is embedded in the downsampling stage of each layer of the dual-stream network. This unit performs max pooling and average pooling along the channel dimension on the thermal stream feature map to generate a thermally saliency attention map, which is then fused with the physical prior mask to form a spatial weight matrix, which is multiplied pointwise with the RGB stream feature map. This operation forces the network to focus on structural damage areas with real thermal anomalies at the feature level, significantly improving the feature response intensity for fine cracks and deep voids in dim or occluded environments.
[0054] Finally, the system performs multi-scale decoding and damage quantification assessment. The interactively processed dual-stream features are fused and upsampled at the decoder, outputting semantic segmentation maps corresponding to cracks, voids, and leaks. The system combines the segmentation results to calculate the planar geometric parameters of the damage and extracts the maximum temperature gradient (ΔTmax) within the damaged area. Using a concrete heat conduction inversion model, the system calculates the depth risk level of voids or cracks, ultimately generating a comprehensive structural durability score that includes geometric quantification indicators and depth assessment.
[0055] Example 3 like Figure 5 As shown, this embodiment details the internal architecture and data processing flow of a multimodal analysis module deployed in the cloud or a high-performance workstation. Addressing the pain points in industrial building inspection, such as heterogeneous data modalities, difficulties in qualitative analysis of single information sources, and the susceptibility of general large models to illusions, a deep learning pipeline including feature encoding, semantic alignment, knowledge retrieval, and joint reasoning is constructed. The specific process is as follows.
[0056] First, in the vectorization and encoding stage of multi-source heterogeneous data, a dedicated encoder is used to map the heterogeneous sensor data to a unified high-dimensional feature space. Specifically, for visible light images and infrared thermal images, a Visual Transformer (ViT) is used as the backbone network. The image is segmented into fixed patches and input into the Transformer layer through linear projection and position encoding to extract visual feature vectors (Evis) representing surface texture and internal heat distribution. For the geometric parameters obtained by the line laser scanning unit, a PointNet++ network is used to effectively extract local geometric feature vectors (Egeo) such as crack depth, width change rate, and surface roughness through multi-scale grouping and ensemble abstraction layers. For the physical hardness parameters collected by the digital rebound unit and the in-situ ion concentration titration data collected by the ion concentration detection unit, after normalization processing, the data is input into a multilayer perceptron (MLP) and mapped to physicochemical feature vectors (Ephy) in a high-dimensional latent space through fully connected layers.
[0057] Secondly, a cross-attention mechanism is introduced to perform semantic space alignment of multimodal features. Given that the aforementioned feature vectors reside in separate manifold spaces and cannot be directly concatenated, the visual feature vector (Evis) with the highest information density is selected to construct the query vector (Query, Q). The geometric feature vector (Egeo) and the structured data feature vector (Ephy) are mapped to key (Key, K) and value (Value, V) vectors. An attention weight matrix is generated by calculating the similarity matrix between Q and K, quantifying the correlation strength between visual features and geometric depth and physical hardness. The weighted fused features are then processed through residual connections and a linear projection layer to obtain a fused embedding vector. This fused embedding vector is then mapped to a text embedding space adapted to the Large Language Model (LLM) to generate multimodal features.
[0058] Furthermore, a Retrieval Enhanced Generation (RAG) architecture is employed for compliance reasoning to suppress model illusions and ensure the engineering rigor of diagnostic conclusions. A vectorized professional knowledge base is pre-built, containing the *Code for Design of Durability of Concrete Structures* (GB / T50476), the *Standard for Durability Assessment of Existing Concrete Structures* (GB / T51355-2019), corrosion mechanism literature, and historical typical cases. During the reasoning phase, Top-K semantic retrieval is performed in the knowledge base using the aforementioned generated multimodal fusion tokens, assembling the most relevant code provisions (such as chloride ion content limits) and similar historical cases into contextual hints. Subsequently, the multimodal features and contextual hints are input into a large language model fine-tuned using low-rank adaptation (LoRA) technology. The large language model combines the perceived physical characteristics of damage with the retrieved code constraints to perform logical deduction, outputting a structured assessment report containing damage level assessment, causal analysis, and maintenance recommendations.
[0059] Finally, a data feedback mechanism is established, using actual engineering measurement data as input and constructing a continuous learning loop through low-rank adaptation (LoRA) fine-tuning technology to adapt to the differentiated characteristics of specific industrial scenarios. The mechanism uses multimodal data collected by the wall-climbing detection module as input and the final report and evaluation results from laboratory analysis as ground truth labels. By freezing the main weights of the pre-trained large language model and updating only the parameters of the bypass low-rank matrix, parameter iteration can be completed with extremely low computational resource consumption, continuously improving the diagnostic accuracy and robustness of the large language model in specific factory environments.
[0060] The air-ground collaborative multimodal industrial building durability testing system and method of the present invention have the following other advantages: This invention employs a processing unit to identify defects in building structure surface image data, achieving an overall average accuracy of 89.0%. The accuracy of identifying appearance defects such as cracks, spalling, exposed reinforcement, and leakage has been effectively improved. In terms of comprehensiveness, the missed detection rate for cracks is 17.2%, the missed detection rate for spalling is 8.9%, and the overall average missed detection rate is 14.73%. These indicators are all at the industry-leading level under the harsh conditions of high temperature, high humidity, strong corrosion, and high dust in industrial buildings, providing highly reliable detection support for structural safety management.
[0061] It is understood that this invention has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of this invention. Furthermore, under the teachings of this invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of this invention. Therefore, this invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this invention are within the protection scope of this invention.
Claims
1. A multimodal industrial building durability testing system with air-ground collaboration, characterized in that, include: The air-ground collaborative perception module is used to perform 3D reconstruction of the building interior and simultaneously collect image data of the building structure surface. The 3D reconstruction is used to obtain a 3D environment map of the building interior. The image data of the building structure surface is used to identify defects and obtain the planar position of the suspected defect area of the building structure. Combining the 3D environment map and the real-time pose information when the images were collected, the planar position of the suspected defect area is mapped to the 3D environment map coordinate system to determine the 3D spatial coordinates of the suspected defect area and generate a detection task list containing the 3D coordinates of the suspected defect area. The wall-climbing detection module is used to receive the detection task list, autonomously navigate to the target area based on the three-dimensional spatial coordinates of the suspected defect area in the detection task list, and perform close-range operations. It operates in an orderly manner according to the progressive process of surface cleaning, multi-dimensional physical detection, and micro-destructive physical sampling, and simultaneously collects visible light images, infrared thermograms, physical hardness parameters, geometric parameters and in-situ ion concentration titration data of the target area. The multimodal analysis module is used to receive 3D environmental maps, building structure surface image data, planar location and 3D spatial coordinates of suspected defect areas, visible light images, infrared thermograms, physical hardness parameters, geometric parameters and in-situ ion concentration titration data, and process and analyze them to obtain an industrial building structure durability assessment report that includes damage type, degree, distribution and durability level.
2. The air-ground collaborative multimodal industrial building durability testing system according to claim 1, characterized in that, The air-ground collaborative perception module includes: a drone for autonomous flight inside the building and acquisition of real-time pose information during its flight; a lidar unit mounted on the drone for collecting environmental point cloud data inside the building; an industrial camera mounted on the drone for synchronously acquiring image data of the building's structural surface; and a processing unit located inside the drone for receiving environmental point cloud data inside the building and image data of the building's structural surface, processing the environmental point cloud data inside the building to generate laser SLAM front-end odometry information, using the laser SLAM front-end odometry information to correct the drone's real-time pose information, completing the 3D reconstruction of the building's interior, generating a globally consistent 3D environmental map, and simultaneously identifying defects in the image data of the building's structural surface, obtaining the planar positions of suspected defect areas of the building structure, mapping the planar positions of the suspected defect areas to the 3D environmental map coordinate system, determining the 3D spatial coordinates of the suspected defect areas, and generating a detection task list containing the 3D coordinates of the suspected defect areas.
3. The air-ground collaborative multimodal industrial building durability testing system according to claim 2, characterized in that, When the processing unit processes the environmental point cloud data inside the building, it sequentially performs timestamp synchronization, motion distortion removal, point cloud registration, and cumulative error elimination on the environmental point cloud data inside the building, and calculates the relative pose between frames in real time during the point cloud registration process.
4. The air-ground collaborative multimodal industrial building durability testing system according to claim 2, characterized in that, The air-ground collaborative sensing module also includes a visualization terminal, which receives a three-dimensional environment map. Based on historical maintenance records, original design drawings, and on-site intuitive judgment, engineers mark the three-dimensional coordinates of specific areas of interest on the three-dimensional environment map and send them to the processing unit. The processing unit merges the three-dimensional coordinates of suspected defect areas with the three-dimensional coordinates of specific areas of interest, and prioritizes them based on the importance of the areas and the confidence of the defects, generating a detection task list that includes the priority ranking.
5. The air-ground collaborative multimodal industrial building durability testing system according to claim 2, characterized in that, The air-ground cooperative perception module also includes a hierarchical navigation planning unit, which is installed inside the UAV. The hierarchical navigation planning unit is used to fuse the laser SLAM front-end odometry information with the real-time pose information of the UAV after correction using the laser SLAM front-end odometry information to obtain the real-time pose information of the image data of the building structure surface, and combine it with the three-dimensional environment map to generate the optimal navigation trajectory for the UAV that meets the kinematic constraints.
6. The air-ground collaborative multimodal industrial building durability testing system according to claim 1, characterized in that, The wall-climbing detection module includes: a negative pressure adsorption walking unit for walking on the surface of the building structure; a cleaning unit for cleaning the target area; a dual-spectrum visual acquisition unit for acquiring visible light images and infrared thermal images of the target area; a digital rebound unit for acquiring physical hardness parameters of the target area; a line laser scanning unit for acquiring geometric parameters of the target area; an ion concentration detection unit for acquiring in-situ ion concentration titration data of the target area; a micro-destructive physical sampling unit for performing micro-destructive physical sampling of the target area; and a control unit for receiving the detection task list, controlling the negative pressure adsorption walking unit to autonomously navigate to the target area to perform close-range operations based on the three-dimensional spatial coordinates of suspected defect areas in the detection task list, and controlling the cleaning unit, dual-spectrum visual acquisition unit, digital rebound unit, line laser scanning unit, ion concentration detection unit, and micro-destructive physical sampling unit to operate in an orderly manner according to a progressive process.
7. The air-ground collaborative multimodal industrial building durability testing system according to claim 1, characterized in that, The multimodal analysis module includes a data processing unit, used to receive 3D environmental maps, building structure surface image data, planar positions and 3D spatial coordinates of suspected defect areas, visible light images, infrared thermograms, physical hardness parameters, geometric parameters, and in-situ ion concentration titration data. It extracts texture features from the building structure surface image data and visible light images, temperature field features from the infrared thermograms, geometric topological features from the geometric parameters, and numerical sequence features from the physical hardness parameters and in-situ ion concentration titration data. Then, through a cross-attention mechanism, it maps the features of different modalities to a unified high-dimensional semantic space for alignment and fusion. Combining the 3D environmental map and the planar positions and 3D spatial coordinates of suspected defect areas, it constructs a multimodal embedding vector that comprehensively represents the health status of the building structure. The decision-making unit is used to pre-build a knowledge base containing building structure durability design specifications, industrial corrosion mechanism literature, and historical testing cases. A large language model is trained using these knowledge bases. Multimodal embedding vectors are input into the large language model, which then searches and retrieves the most relevant standard provisions or similar historical testing cases from the knowledge base in real time. These are used as contextual prompts to provide the large language model with standard bases and case references for durability assessment. This assists the large language model in conducting comprehensive reasoning analysis, completing the analysis and diagnosis of the durability status of industrial buildings, and generating a durability assessment report for industrial building structures that includes damage type, degree, distribution, and durability level.
8. The air-ground collaborative multimodal industrial building durability testing system according to claim 7, characterized in that, When the data processing unit extracts the texture features of the building structure surface image data and visible light image, the temperature field features of the infrared thermogram, the geometric topological features of the geometric parameters, and the numerical sequence features of the physical hardness parameters and in-situ ion concentration titration data, it uses a visual converter to extract the texture features of the building structure surface image data and visible light image, uses a point cloud network++ to extract the geometric topological features of the geometric parameters, and uses a multilayer perceptron to extract the temperature field features of the infrared thermogram and the numerical sequence features of the physical hardness parameters and in-situ ion concentration titration data.
9. The air-ground collaborative multimodal industrial building durability testing system according to claim 7, characterized in that, When the data processing unit maps features from different modalities to a unified high-dimensional semantic space for alignment and fusion through a cross-attention mechanism, it employs a multi-source heterogeneous feature alignment technique.
10. A method for testing the durability of multimodal industrial buildings using a combined air-ground approach, characterized in that, The system described in any one of claims 1 to 9 comprises the following steps: The interior of the building is reconstructed in 3D and the surface image data of the building structure are collected simultaneously. The 3D reconstruction is used to obtain a 3D environment map of the interior of the building. The surface image data of the building structure is used to identify defects and obtain the planar position of the suspected defect area of the building structure. Combining the 3D environment map and the real-time pose information when the images are collected, the planar position of the suspected defect area is mapped to the coordinate system of the 3D environment map to determine the 3D spatial coordinates of the suspected defect area and generate a detection task list containing the 3D coordinates of the suspected defect area. Based on the three-dimensional spatial coordinates of the suspected defective areas in the inspection task list, the system autonomously navigates to the target area to perform close-range operations. It operates in an orderly manner according to the progressive process of surface cleaning, multi-dimensional physical detection, and micro-destructive physical sampling, while simultaneously collecting visible light images, infrared thermal images, physical hardness parameters, geometric parameters, and in-situ ion concentration titration data of the target area. The process involves processing and analyzing 3D environmental maps, building structure surface image data, planar locations and 3D spatial coordinates of suspected defect areas, visible light images, infrared thermograms, physical hardness parameters, geometric parameters, and in-situ ion concentration titration data to obtain an industrial building structure durability assessment report that includes damage type, degree, distribution, and durability level.