A structural crack identification and positioning method based on multi-modal deep learning

CN121706007BActive Publication Date: 2026-08-21XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511900452.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-08-21
Estimated Expiration
2045-12-16

AI Technical Summary

Technical Problem

[0004]针对现有技术的不足,本发明提供了一种基于多模态深度学习的结构裂缝识别与定位方法,解决了无法细致确定裂缝具体位置、几何特征,最终导致结构裂缝的识别与定位整体不准确的问题

Benefits of technology

(1)、该基于多模态深度学习的结构裂缝识别与定位方法,通过多任务学习架构,充分发挥卷积神经网络对振动、视觉多模态数据的特征提取能力,可同时完成裂缝“存在性分类”与“位置坐标回归”,实现显性裂缝的端到端高效检测,既提升了识别精度,又简化了检测流程,避免多步骤单独处理带来的误差累积。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121706007B_ABST
    Figure CN121706007B_ABST
Patent Text Reader

Abstract

The application discloses a kind of structural crack identification and positioning method based on multi-modal deep learning, it is related to crack identification and positioning technical field.The steps of the method include: collecting building structure prior information, generating controllable micro-vibration excitation sequence and exciting, vibration signal and image data are synchronously collected;According to vibration signal, image data is marked respectively abnormal area, after correlation analysis, crack candidate area is screened;Candidate area data is extracted, and convolutional neural network explicit crack identification model is constructed, and the result containing explicit crack area is output by input data;Based on explicit result, feature template library is established, and the feature of non-explicit candidate area is extracted and compared with template library, and the identification result of hidden crack is obtained by combining infrared verification result;Integrate explicit and hidden crack result, generate structure crack detection report.Based on the structure crack detection report, structural crack identification and positioning are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of crack identification and localization engineering, specifically a structural crack identification and localization method based on multimodal deep learning. Background Technology

[0002] As the service life of buildings increases and external loads and environmental conditions take effect over a long period of time, structural cracking has gradually become a key factor threatening the durability and safety of buildings.

[0003] Currently, structural crack detection is achieved using vibration signal-based methods. By applying controllable micro-vibration excitation based on structural zoning design to a building, the presence of cracks alters the stiffness distribution of components, leading to changes in vibration transmission paths and energy dissipation characteristics. By capturing these differences in mechanical response, crack regions can be located. However, while vibration signal-based detection methods can capture abnormal responses under structural stress, the lack of morphological information makes it difficult to accurately determine whether an anomaly is a crack, nor can it precisely pinpoint the crack's location and geometric features, ultimately resulting in overall inaccurate identification and location of structural cracks. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a structural crack identification and localization method based on multimodal deep learning, which solves the problem that the specific location and geometric features of cracks cannot be precisely determined, ultimately leading to inaccurate overall identification and localization of structural cracks.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for structural crack identification and localization based on multimodal deep learning, comprising the following steps: Step S1: Collect prior information about the building structure, generate a controllable micro-vibration excitation sequence based on the prior information about the building structure, apply slight vibration to the building according to the controllable micro-vibration excitation sequence, and simultaneously collect the vibration signal and image data of the building; Step S2: Based on the vibration signal, mark the vibration anomaly area; based on the image data, mark the image anomaly area; by performing correlation analysis between the vibration anomaly area and the image anomaly area, select candidate crack areas; Step S3: Extract vibration signals and image data of the crack candidate region, construct a visible crack identification model based on a convolutional neural network, input the vibration signals and image data of the crack candidate region into the visible crack identification model, and output the visible crack identification result, which includes the visible crack region. Step S4: Based on the results of the visible crack identification, construct a visible crack feature template library, extract features from candidate regions that are not identified as visible cracks, generate a candidate region feature vector set, compare the candidate region feature vector set with the visible crack feature template library to obtain feature differences, verify the candidate regions that are not identified as visible cracks using infrared technology to obtain infrared thermal feature verification results, and combine the feature differences and infrared thermal feature verification results to obtain the hidden crack identification results. Step S5: Based on the results of the identification of visible cracks and the results of the identification of hidden cracks, generate a structural crack detection report to realize the identification and location of structural cracks.

[0006] Preferably, the process of collecting prior information about the building structure and generating a controllable micro-vibration excitation sequence based on that prior information includes: The prior information of the target building structure was collected as follows: First, collect the design documents, as-built reports, and operation and maintenance records of the target building and organize them into a basic information database; The controllable micro-vibration excitation sequence based on structural partitioning is designed as follows: First, the excitation parameter constraints are determined for each sub-region. Based on its mechanical parameters Calculate the range of excitation parameter constraints: The frequency constraint is estimated using the formula for the natural frequency of the component:

[0007] in The span of the component is expressed in meters (m). This refers to the material density, in units of... ; The cross-sectional area is expressed in meters (m²). 2 ; I is the elastic modulus of the material, measured in Pa; I is the moment of inertia of the cross section, measured in meters. 4 Determine the excitation frequency range , usually take ; Amplitude constraints are used to determine the maximum excitation force through preliminary experiments. Ensure structural strain ; Duration constraint is the duration of single-region excitation. ; Based on the above-mentioned excitation parameter constraints, for each sub-region Generate a dedicated incentive sequence .

[0008] Preferably, applying slight vibrations to the building according to the controllable micro-vibration excitation sequence, and simultaneously collecting the building's vibration signals and image data includes: According to the excitation sequence Vibrations were applied sequentially to each sub-region, and the data was recorded synchronously as follows: Vibration sensors collect acceleration signals Stored sequentially in .csv format; Camera captures continuous image sequences Store as a .png file, along with the associated sub-region ID and excitation parameters.

[0009] Preferably, based on the vibration signal, abnormal vibration areas are marked; based on the image data, abnormal image areas are marked including: The correlation analysis algorithm is used to identify regions that are in phase with the anchor point vibration but have abnormal intensity. The process of marking the vibration anomaly region is as follows: For each monitoring point, the preprocessed vibration acceleration signal is extracted. sequence Simultaneously, extract the preprocessed vibration acceleration signal sequence of the anchor point. Calculate the Pearson correlation coefficient between the two. :

[0010] in, It is the preprocessed vibration acceleration signal from the monitoring point. In the Each sampling time The value of , The vibration acceleration signal of the anchor point after preprocessing is at the 1st Each sampling time The value of , This represents the number of signal sampling points. It is the average value of the signals from the monitoring points. It is the mean of the anchor signal; the numerator is the covariance of the two signals, and the denominator is the product of the standard deviations of the two signals; like The vibration in this area was determined to be in phase with that of the anchor point, and it was marked as an abnormal vibration area. Its spatial coordinate range was recorded. ; Based on preprocessed normalized image sequences The process of marking abnormal regions in an image using threshold analysis is as follows: For each frame of the image, contours are extracted using Canny edge detection, retaining continuous edges with a length > 10 pixels. For a given edge detected in the k-th frame... Suppose it contains Pixels, coordinates are Then the coordinates of the center point of the edge are :

[0011] in, It is the first coordinates of pixels in These are horizontal pixel coordinates. These are vertical pixel coordinates; It is the first frame edges The coordinates of the horizontal center point, through the... The average of the horizontal coordinates of each pixel is obtained; It is the first frame edges The coordinates of the vertical center point, through the... The average of the vertical coordinates of each pixel is used to obtain the value, where m is the number of pixels. For the same edge matched in two consecutive frames, calculate the center point coordinates respectively. and The Euclidean distance between the two is:

[0012] in, It is the Euclidean distance between the center points of the edges of two frames. These are the coordinates of the center point of the edge in the (k+1)th frame; The timing stability is determined as follows: For 20 consecutive frames of images, calculate the displacement between every two adjacent frames. Take the average displacement. k is the frame number index; like Pixels are identified as temporally stable edges and marked as image anomaly regions. The pixel coordinates of these anomaly regions are then converted to 3D architectural coordinates via spatial mapping and recorded. .

[0013] Preferably, vibration signals and image data of the crack candidate region are extracted, and a visible crack identification model is constructed based on a convolutional neural network. The vibration signals and image data of the crack candidate region are input into the visible crack identification model, and the output visible crack identification results include: Based on convolutional neural networks, a visible crack identification model is constructed as follows: The input layer is divided into two branches; One branch is the vibration input. For the crack candidate region output in step S2, the acceleration signal of the corresponding vibration sensor is extracted, and high-frequency noise is removed by a 50Hz low-pass filter to obtain the preprocessed signal. ; Another branch uses visual input, mapping the 3D coordinates of the crack candidate region to the image plane and cropping the ROI region. Enhance crack edges through adaptive threshold segmentation; The feature extraction branch is divided into a vibration feature extraction branch and a visual feature extraction branch. The bidirectional attention fusion layer is as follows: First, the correlation matrix is ​​calculated to generate the cross-correlation matrix of vibration-visual features. :

[0014] Where i and j are the indices of the vibration dimension and the visual dimension, respectively, quantifying the correlation between the i-th vibration dimension and the j-th visual dimension; Weighted fusion is performed, with vibration and visual features weighted separately and then stitched together to output a 768-dimensional fused feature. :

[0015] in, It is the weight vector of vibration characteristics. The weight vector of vibration characteristics, Element-wise multiplication For dimensional splicing, It is the vibration characteristic vector. It is a visual feature vector; By using a shared feature layer, the features output from the vibration feature extraction branch and the visual feature extraction branch are further fused, compressed, and abstracted to obtain more representative deep features. ; The final output layer is a multi-task output head, which is based on the deep representation output by the shared feature layer. It simultaneously performs two tasks: determining the existence of cracks and locating cracks, in order to meet the actual needs of crack detection. It is divided into a classification head and a regression head, and the classification head is output first, followed by the regression head. The classification header output process is as follows: This part uses a fully connected layer to represent the 256-dimensional depth. The output is converted to one dimension, and then the Sigmoid activation function is used to map the output value to the interval [0, 1] to obtain the probability of crack existence. , The range of values ​​for is [0, 1]. The closer the probability is to 1, the more likely the input region is to be a visible crack; the model takes the output probability. The candidate cracks were determined to be explicit cracks; The regression head output process is as follows: Using fully connected layers, the classification head determines a 256-dimensional depth representation of the dominant split. Convert to 4D output, outputting the coordinates of the explicit crack boundary. .

[0016] Preferably, based on the results of visible crack identification, a visible crack feature template library is constructed, including: Based on the visible cracks identified in step S3, their multi-dimensional features are extracted to construct a visible crack feature template library, providing a benchmark reference for the detection of hidden cracks. For depth characterization extraction, the depth characterization corresponding to the visible cracks is extracted from the multi-task network in step S3. ; For vibration and visual feature extraction, vibration features of visible cracks are extracted simultaneously. Visual features ; Infrared feature extraction was performed to acquire infrared thermal images of the visible crack region and obtain the temperature gradient. (Temperature gradient reflects the degree of spatial variation in the temperature field of the crack region), the average temperature difference between the crack and the background. , Quantify the thermal difference between the crack region and the normal background. It is the average temperature of the crack region. It is the average temperature of the background area; Organize the above features into a template library according to crack number. .

[0017] Preferably, feature extraction is performed on candidate regions not identified as overt cracks to generate a candidate region feature vector set. The candidate region feature vector set is then compared with the overt crack feature template library to obtain feature differences, including: For the candidate regions output in step S2 that were not identified as explicit cracks in step S3, the feature extraction process in step S3 is repeated to obtain the following multimodal features: For vibration features, vibration signals from candidate regions are extracted and processed using a 1D-CNN. ; For visual features, the image is cropped from the candidate region and then processed through a newly added convolutional layer in VGG16+ to obtain... ; Deep representation, will and Inputting the bidirectional attention and shared feature layer from step S3 yields a deep representation of the candidate region. ; By integrating the above data, a set of feature vectors for candidate regions is obtained. ); By comparing the generated candidate region feature vector set with the explicit crack feature template library, feature differences are obtained. These feature differences include depth differences, vibration differences, and visual differences, as follows: For depth differences, calculate the depth representation vector of the candidate region. The deep representation template vector most similar to the template library Euclidean distance between :

[0018] Among them, the Euclidean distance reflects the straight-line distance between two vectors in space. The smaller the distance, the more similar the depth representation of the candidate region is to the depth representation of the explicit crack, and the more likely the candidate region is to be a hidden crack. For vibration differences, calculate the vibration feature vector of the candidate region. Vibration feature template vector most similar to the template library The cosine similarity is calculated by subtracting 1 from the cosine similarity. :

[0019] in, It refers to vibration differences. Cosine similarity measures the cosine of the angle between two vectors. The smaller the angle, the higher the cosine similarity, and the more matched the vibration characteristics. The smaller the value, the higher the vibration characteristic matching degree; For visual differences, similar to the calculation of vibration differences, the visual feature vectors of the candidate regions are calculated. The visual feature template vector most similar to the template library The cosine similarity is calculated, and 1 is subtracted from the cosine similarity to obtain the result. :

[0020] in, Visual differences The smaller the value, the higher the visual feature matching degree.

[0021] Preferably, candidate areas not identified as visible cracks are verified using infrared technology, and the infrared thermal feature verification results include: Candidate regions not identified as visible cracks were verified using infrared technology, and the infrared thermal feature verification results were obtained: Candidate regions not identified as visible cracks were verified using infrared technology, and the infrared thermal feature verification results were obtained: Temperature gradient direction verification checks the temperature gradient direction of the candidate region against the temperature gradient template vector most similar to the template vector in the template library. If the angle between the directions represented is less than or equal to 15°, it means that the heat conduction direction of the crack is relatively consistent with the direction of the visible crack, which supports the judgment that the candidate area is a hidden crack. Temperature difference verification is used to determine the average temperature difference of the candidate region. Is the temperature greater than or equal to 0.6°C? If so, it indicates that there is a change in thermal resistance in the candidate region due to the crack, providing evidence from a thermal perspective that the candidate region is a hidden crack.

[0022] Preferably, the hidden crack identification result obtained by combining the aforementioned feature differences and infrared thermal feature verification results includes: A candidate area will be identified as a hidden crack only if all three of the following conditions are met: The first condition is that the depth difference meets the standard, which is the depth difference between the candidate region and the visible crack template. It must be greater than or equal to the difference in average depth between visible cracks. 25% of ; The second condition is mode consistency and vibration characteristic differences. The difference in visual features must be less than or equal to 0.4. It must also be less than or equal to 0.4, that is... and ; The third condition is that the infrared verification is passed, which is one of the conditions for infrared verification. Infrared verification is divided into temperature gradient direction verification and temperature difference verification. Temperature gradient direction verification checks the candidate region's temperature gradient direction against the most similar temperature gradient template vector in the template library. If the angle between the directions represented is less than or equal to 15°, it means that the heat conduction direction of the crack is relatively consistent with the direction of the visible crack, which supports the judgment that the candidate area is a hidden crack. Temperature difference verification is used to determine the average temperature difference of the candidate region. If the temperature is greater than or equal to 0.6°C, it indicates that there is a change in thermal resistance in the candidate region due to the crack, providing evidence from a thermal perspective that the candidate region is a hidden crack.

[0023] Candidate areas that meet the above criteria will be marked as hidden cracks.

[0024] Preferably, based on the identified visible cracks and the identified hidden cracks, the structural crack detection report is generated including: The identification and location results of cracks are clearly presented to relevant personnel through intuitive visualization and detailed reports. The process is as follows: Using structural BIM, crack markers are overlaid on the model. Visible cracks are marked with solid red lines, and the line width increases with the damage index. The width of the line increases with the increase of the damage index DI, which makes it easy to visually distinguish visible cracks of different damage levels; for hidden cracks, a yellow dashed line is used to mark them, and the line width increases with the increase of the damage index DI, which makes it easy to quickly identify the condition of hidden cracks. The identification and location report is as follows: For the crack statistics section, a detailed count of the total number of cracks, as well as the proportion of visible cracks and hidden cracks, is provided. The distribution of cracks of different grades is also listed to give personnel a macroscopic understanding of the overall crack situation. For the location details section, the three-dimensional coordinates of each crack are precisely listed. Combined with morphological feature information, the specific location and extension direction of the crack in the structure are clearly described, providing accurate location information for subsequent maintenance work.

[0025] This invention provides a method for structural crack identification and localization based on multimodal deep learning, involving machine learning and deep learning technologies, which has the following beneficial effects: (1) The structural crack identification and localization method based on multimodal deep learning fully leverages the feature extraction capabilities of convolutional neural networks for vibration and visual multimodal data through a multi-task learning architecture. It can simultaneously complete the "existence classification" and "location coordinate regression" of cracks, achieving end-to-end efficient detection of visible cracks. This not only improves the identification accuracy but also simplifies the detection process and avoids the accumulation of errors caused by processing multiple steps separately.

[0026] (2) This structural crack identification and localization method based on multimodal deep learning dynamically mines and enhances the correlation features between the two modes related to cracks by calculating the cross-correlation matrix of vibration features and visual features, while suppressing irrelevant noise interference. It can accurately capture the most critical information for crack identification in multimodal data, significantly improve the pertinence and accuracy of multimodal feature fusion, and thus improve the accuracy of crack identification.

[0027] (3) A structural crack identification and localization method based on multimodal deep learning introduces physical features such as temperature gradient and temperature difference from infrared thermal imaging to perform cross-modal physical-level verification of high-suspicious hidden cracks. By utilizing the temperature difference between cracked and healthy areas, the uncertainty of vibration and visual modalities in the identification of hidden cracks is compensated, significantly reducing the misjudgment rate of hidden cracks and improving the reliability of its judgment.

[0028] (4) A structural crack identification and localization method based on multimodal deep learning constructs a multimodal feature template library of visible cracks based on the identification results of real visible cracks. This library serves as the core reference benchmark for counterfactual comparison. By performing cross-modal difference analysis with the feature representations of candidate regions not identified as visible cracks, the deviation between the features of candidate regions and typical crack features can be accurately quantified. This counterfactual comparison logic, which uses real crack features as a reference, can effectively filter out non-crack interference signals, accurately anchor highly suspicious areas of hidden cracks, and significantly improve the targeting and preliminary screening efficiency of hidden crack identification. Attached Figure Description

[0029] Figure 1 This is a flowchart of a structural crack identification and localization method based on multimodal deep learning proposed in this invention; Figure 2 This is a hierarchical diagram of the explicit crack identification results obtained by the structural crack identification and localization method based on multimodal deep learning proposed in this invention. Figure 3 This is a hierarchical diagram of the hidden crack identification results obtained in the structural crack identification and localization method based on multimodal deep learning proposed in this invention. Detailed Implementation

[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0031] Please see Figure 1-3 This invention provides a technical solution: a method for structural crack identification and localization based on multimodal deep learning. Specifically, the method for structural crack identification and localization based on multimodal deep learning is provided below. Figure 1 The method includes the following steps: Step S1: Collect prior information about the building structure, generate a controllable micro-vibration excitation sequence based on the prior information about the building structure, apply slight vibration to the building according to the controllable micro-vibration excitation sequence, and simultaneously collect the vibration signal and image data of the building; The prior information of the target building structure was collected as follows: First, collect the target building's design documents (structural construction drawings, component reinforcement drawings, material parameter tables), completion reports (including structural joint acceptance records), and operation and maintenance archives (historical crack detection records and maintenance records), and organize them into a basic information database.

[0032] By conducting an on-site inspection of the building and combining this information with the aforementioned database, the following information was verified: For the overall structural layout, confirm the floor division, functional zoning (such as load-bearing area and non-load-bearing area) and component distribution (actual location and size of beams, columns and slabs). Regarding the mechanical parameters of the components, based on the material parameter tables in the design documents (such as concrete strength grade C30, steel reinforcement type HRB400) and the material inspection records in the completion report, the relevant material specifications are consulted, and on-site inspections are conducted to jointly determine the elastic modulus E and material density of the components. Based on the cross-sectional dimensions in the structural construction drawings and component reinforcement drawings, determine the cross-sectional area A of the component; For structural joint parameters, mark the spatial coordinates of preset joints such as settlement joints and expansion joints. (3D coordinates), orientation angle, and width, recorded as a set. ( (Total number of structural joints) The detection sub-regions are divided based on a two-dimensional approach: force characteristics and spatial distribution. Divide the building vertically into large areas (e.g., ... (Number of floors); each large area is horizontally subdivided into sub-areas according to component type (beam system, column group, floor slab). (Total number of sub-regions), and define the boundary coordinates of each sub-region. For each sub-region The associated components are modeled using mechanical parametric methods to form their set of mechanical parameters. ; Output structured partitioning model ,in The set of structural seams within the sub-region (including structural seam parameters); The controllable micro-vibration excitation sequence based on structural partitioning is designed as follows: First, the excitation parameter constraints are determined for each sub-region. The range of excitation parameter constraints is calculated based on its mechanical parameters. These constraints include frequency constraints, amplitude constraints, and duration constraints. The frequency constraint is calculated using the formula for estimating the natural frequency of a component:

[0033] in The span of the component is expressed in meters (m). For material density, ; The cross-sectional area is expressed in meters (m²). 2 ; I is the elastic modulus of the material, measured in Pa; I is the moment of inertia of the cross section, measured in meters. 4Determine the excitation frequency range , usually take ; The amplitude constraint is determined through preliminary experiments to ensure that the structural strain ε≤100με is within the elastic range. The corresponding excitation force (such as the force amplitude of the vibrator) needs to be determined by calculation and preliminary experiments based on the mass, stiffness and required response level of the target component. Duration constraint is the duration of single-region excitation. (Including a 3s frequency sweep phase + a 5~9s frequency fixation phase).

[0034] Based on the above-mentioned excitation parameter constraints, for each sub-region Generate a dedicated incentive sequence And it meets the following conditions: First, select the excitation point, avoiding structural seams. Preferably located at potential stress concentration locations within the sub-region (such as beam ends and column corners), denoted as coordinates.

[0035] The frequency mode uses a linear sweep frequency for the first 3 seconds. ),back Fixed-frequency excitation was used (the frequency with the most significant vibration response during the frequency sweep process was selected). ; The excitation sequence is from bottom to top, from non-critical areas to critical areas, with an excitation interval of ≥30s between adjacent areas (to avoid vibration superposition interference).

[0036] For the vibration monitoring module, a triaxial accelerometer with an accuracy ≥0.01g is used in each sub-region. Four measuring points were set up (covering both ends and the middle of the component), and three anchor point sensors were set up (placed in the verified crack-free healthy area), with the sampling rate set to 1000Hz; For the image acquisition module, an industrial camera with a resolution of ≥5 million pixels (equipped with an 8mm fixed-focus lens and an adjustable fill light) is used. One to two cameras are deployed in each sub-area to ensure that the lens field of view covers the entire sub-area. The frame rate is set to 100fps (to match the vibration signal sampling rate for timing analysis). To ensure time synchronization, a GPS timing module (synchronization error ≤ 1ms) is used to provide a unified timestamp T for all devices, ensuring that the time dimensions of vibration signals, image frames, and environmental data are aligned.

[0037] According to the excitation sequence Vibrations were applied sequentially to each sub-region, and the data was recorded synchronously as follows: Vibration sensors collect acceleration signals Stored continuously in .csv format, with timestamps associated. and sub-region ID; Camera captures continuous image sequences ( For frame number, Corresponding time , (For the stimulus start time), store it in .png format, and associate it with the sub-region ID and stimulus parameters; Vibration signal preprocessing is as follows: First, noise filtering is performed using a 4th-order Butterworth low-pass filter (cutoff frequency 500Hz) to remove high-frequency noise, and wavelet threshold denoising is used to process non-stationary interference; then baseline correction is performed on the acceleration signal. The mean is subtracted to eliminate zero drift; finally, the speed is calculated using a quadratic integral. and displacement Extract signal segments within the excitation period and label the corresponding excitation frequencies. .

[0038] Image data preprocessing is as follows: First, distortion correction is performed, and then the image is analyzed based on camera intrinsic parameters. Distortion correction is performed; then enhancement and denoising are applied, using the CLAHE algorithm to enhance local contrast and non-local mean filtering to remove image noise; finally, region cropping is performed based on sub-region boundaries. The image is cropped to retain the effective monitoring area, and a standardized image sequence is output. .

[0039] The preprocessed structured bimodal dataset is as follows:

[0040] in, This is the preprocessed vibration signal. This is a preprocessed, standardized image sequence.

[0041] This step aims to establish a detection benchmark through structured information collection, design targeted micro-vibration excitation to stimulate crack response, and construct a high-quality dual-modal dataset through synchronous acquisition and preprocessing to provide reliable input for subsequent analysis.

[0042] Step S2: Based on the vibration signal, mark the vibration anomaly area; based on the image data, mark the image anomaly area; by performing correlation analysis between the vibration anomaly area and the image anomaly area, select candidate crack areas.

[0043] Correlation analysis algorithms are used to identify regions that are in phase with the vibration of the anchor point (a healthy area without cracks) but have abnormal intensity. The process of marking the vibration anomaly areas is as follows: For each monitoring point, the preprocessed vibration acceleration signal is extracted. sequence ( , (Sampling time); simultaneously extract the preprocessed vibration acceleration signal sequence of the anchor point. Calculate the Pearson correlation coefficient between the two. :

[0044] in, It is the preprocessed vibration acceleration signal from the monitoring point. In the Each sampling time The value of , The vibration acceleration signal of the anchor point after preprocessing is at the 1st Each sampling time The value of , This represents the number of signal sampling points. It is the average value of the signals from the monitoring points. It is the mean of the anchor point signal; the numerator is the covariance of the two signals, and the denominator is the product of the standard deviations of the two signals (serving a standardization function, making...). The range is limited to [-1, 1]).

[0045] like The vibration in this area was determined to be in phase with that of the anchor point (the vibration rhythm in the crack area was consistent with the overall structure), and it was marked as a vibration anomaly area, with its spatial coordinate range recorded. .

[0046] Based on preprocessed normalized image sequences The process of marking abnormal regions in an image using threshold analysis is as follows: For each frame of the image, contours are extracted using Canny edge detection, retaining continuous edges with a length > 10 pixels. For a given edge detected in the k-th frame... Suppose it contains Pixels, coordinates are ( Horizontal pixel coordinates (where the coordinates are vertical pixel coordinates), then the coordinates of the center point of this edge are... :

[0047] in, It is the first coordinates of pixels in These are horizontal pixel coordinates. These are vertical pixel coordinates; It is the first frame edges The coordinates of the horizontal center point, through the... The average of the horizontal coordinates of each pixel is obtained; It is the first frame edges The coordinates of the vertical center point, through the... The value is obtained by averaging the vertical coordinates of the pixels, where m is the number of pixels.

[0048] For two consecutive frames (the first) Frame and the For the same edge matched in the frame, calculate the center point coordinates respectively. and The Euclidean distance (pixel displacement) between the two is:

[0049] in, It is the Euclidean distance between the center points of the edges of two frames. It is the coordinate of the center point of the edge in the (k+1)th frame.

[0050] The timing stability is determined as follows: For 20 consecutive frames of images, calculate the displacement between every two adjacent frames. Take the average displacement. k is the frame number index; like A pixel (i.e., an average inter-frame displacement of less than 3 pixels) is considered temporally stable and marked as an image anomaly region. The pixel coordinates of the image anomaly region are converted to architectural 3D coordinates through spatial mapping and recorded as follows: .

[0051] Only when the vibration abnormal area With image anomaly area A region is included as an initial crack candidate region only if it meets all of the following conditions: Calculate the vibration anomaly zone With image anomaly area The crossover ratio is given by the formula:

[0052] in, It is the intersection-exchange ratio (IoE) of the vibration anomaly region and the image anomaly region, provided that the following conditions are met. ; Matching is performed using timestamps to obtain the time difference that satisfies both vibration anomaly and image anomaly. (i.e., within 10 frames of the image).

[0053] Finally, by calling the structural seam / window seam database collected in step S1, the Euclidean distance D between the initial crack candidate region and the known interfering seam is calculated. When the distance D... Furthermore, when matching shapes (such as the right-angle features of window seams), this part of the initial crack candidate region is excluded, and the crack candidate region is obtained.

[0054] For the selected crack candidate regions, their boundary coordinates are determined using the minimum bounding rectangle algorithm:

[0055] in, Number the candidate regions. Given the floor height, the final output is a set of candidate crack regions. .

[0056] This step, based on the vibration signal and image sequence output in step S1, identifies abnormal features in the two types of data and establishes association rules to filter out areas that simultaneously meet the requirements of vibration abnormality and image abnormality, eliminates interference from window seams, construction seams, etc., and finally determines the coordinates of the candidate crack area.

[0057] Step S3: Extract vibration signals and image data of the crack candidate region, construct a visible crack identification model based on a convolutional neural network, input the vibration signals and image data of the crack candidate region into the visible crack identification model, and output the visible crack identification result, which includes the visible crack region.

[0058] Based on convolutional neural networks, a visible crack identification model is constructed as follows: The input layer is divided into two branches; One branch is the vibration input. For the crack candidate region output in step S2, the acceleration signal of the corresponding vibration sensor is extracted, and high-frequency noise is removed by a 50Hz low-pass filter to obtain the preprocessed signal. ; Another branch uses visual input, mapping the 3D coordinates of the crack candidate region to the image plane and cropping the ROI region. The crack edges are enhanced through adaptive threshold segmentation.

[0059] The feature extraction branch (dual-modal parallel) is divided into a vibration feature extraction branch and a visual feature extraction branch; The vibration feature extraction branch (1D-CNN) is as follows: This branch focuses on feature extraction from vibration signals, aiming to mine mechanical anomalies related to cracks from continuous vibration data, and is divided into three layers; The first layer, Layer 1, first extracts preliminary features from the input vibration signal through a 1D convolutional layer (equipped with 64 convolutional kernels, kernel size 7, stride 2). Then, BatchNorm is performed to accelerate training and improve stability. Next, ReLU activation function is used to introduce nonlinearity. Finally, max pooling (pooling kernel size 3, stride 2) is used to downsample the features, reducing the number of parameters and retaining key features.

[0060] The second layer, Layer 2, uses a 1D convolutional layer with 128 kernels (size 5, stride 1) to further extract features. Then, BatchNorm and ReLU operations are performed to refine the feature representation of the vibration signal.

[0061] The third layer, Layer 3, uses a 1D convolutional layer with 256 kernels (size 3, stride 1) to deeply mine features. Subsequent BatchNorm, ReLU, and GlobalAvgPool1D operations compress the extracted features into a fixed-length vector, ultimately outputting a 256-dimensional vibration feature vector. This vector implies physical characteristics such as the energy proportion and kurtosis of the vibration signal, which can reflect whether there are changes in the mechanical properties of the structure due to cracks.

[0062] The visual feature extraction branch (2D-CNN) is as follows; This branch focuses on image data, specifically from the Region of Interest (ROI). The morphological features of the cracks were extracted.

[0063] A pre-trained VGG16 model was used, and the parameters of the first 10 layers were frozen. The general features (such as edges and textures) learned by the pre-trained model on large-scale image datasets can provide a good foundation for the recognition of specific targets such as cracks. Freezing the first 10 layers can prevent these general features from being destroyed during training for cracks. By utilizing the first 10 layers of the original VGG16 (containing 5 convolutional blocks and pooling operations), basic features such as low-level edges and textures of cracks in the image are extracted. These are important prerequisites for identifying crack morphology. Three more convolutional layers (512 kernels, size 3) are added, and combined with BatchNorm, RelU, and GlobalAvgPool2D operations, the previously extracted basic features are further abstracted and integrated, finally outputting a 512-dimensional visual feature vector. This vector contains implicit information about the crack's morphological features, such as its length and orientation angle, which is used to subsequently determine whether a visible crack exists in the image region.

[0064] The bidirectional attention fusion layer is as follows: First, the correlation matrix is ​​calculated to generate the cross-correlation matrix of vibration-visual features. :

[0065] Where i and j are the indices of the vibration dimension and the visual dimension, respectively, quantifying the correlation between the i-th vibration dimension and the j-th visual dimension. It is the vibration characteristic vector. It is a visual feature vector.

[0066] Then, attention weights are generated: Vibration characteristic weights (For each vibration feature, its correlation with all visual features is summarized and normalized using Softmax); Visual feature weights (For each visual feature, its correlation with all vibration features is summarized and normalized using Softmax).

[0067] Finally, a weighted fusion process is performed, where vibration and visual features are weighted separately and then concatenated to output a 768-dimensional fused feature. :

[0068] in, It is the weight vector of vibration characteristics. The weight vector of vibration characteristics, Element-wise multiplication Dimensional splicing.

[0069] The shared feature layer (feature compression and abstraction) is as follows: The shared feature layer serves to further fuse, compress, and abstract the features output from the vibration feature extraction branch and the visual feature extraction branch to obtain more representative deep features.

[0070] First, there's a fully connected layer that processes the input 768-dimensional features. The model is converted to 512 dimensions and then subjected to BatchNorm (batch normalization) to accelerate network training and improve stability. Then, non-linearity is introduced through the ReLU activation function to increase the expressiveness of the model. Finally, Dropout (with a dropout rate of 0.3) is used to randomly drop some neurons during training to prevent overfitting.

[0071] Next comes the fully connected layer, which further compresses the 512-dimensional features to 256 dimensions. Then, BatchNorm and ReLU operations are performed to obtain a more concise feature representation that is rich in key information.

[0072] The shared feature layer outputs a 256-dimensional deep representation. This vector integrates key correlation features between vibration and vision, and can comprehensively reflect the relevant information of nonlinear mapping in the input data.

[0073] The final output layer is a multi-task output head, which is based on the deep representation output by the shared feature layer. It simultaneously performs two tasks: determining the existence of cracks and locating their positions, to meet the actual needs of crack detection. It consists of a classification head and a regression head, with the classification head output first, followed by the regression head. The classification header output process is as follows: The classification head (crack existence determination) uses a fully connected layer to represent the 256-dimensional depth. The output is converted to one dimension. Then, the Sigmoid activation function is used to map the output value to the interval [0, 1] to obtain the probability of crack existence. , The range of values ​​for is [0, 1]. The closer the probability is to 1, the more likely the input region is to be a visible crack; the model takes the output probability. The candidate cracks were determined to be visible cracks.

[0074] The regression head output process is as follows: Using fully connected layers, the classification head determines a 256-dimensional depth representation of the dominant split. Convert to 4D output, outputting the coordinates of the explicit crack boundary. The coordinates here are based on image coordinates.

[0075] The model's final output includes the identification results of visible cracks and the coordinates of the visible crack boundaries, which, together with the associated structural member (beam / column) and feature parameters (parameters obtained from feature extraction in the previous steps), form a set. .

[0076] This step accurately identifies the location, shape, and characteristic parameters of visible cracks, not only completing part of the crack detection target, but more importantly, providing core characteristic data (vibration, vision, and depth characterization) for building the "visible crack feature template library" in step S4, which serves as the benchmark for subsequent hidden crack identification.

[0077] Step S4: Based on the results of the visible crack identification, construct a visible crack feature template library, extract features from candidate regions that are not identified as visible cracks, generate a candidate region feature vector set, compare the candidate region feature vector set with the visible crack feature template library to obtain feature differences, verify the candidate regions that are not identified as visible cracks using infrared technology to obtain infrared thermal feature verification results, and combine the feature differences and infrared thermal feature verification results to obtain the hidden crack identification results.

[0078] Based on the explicit cracks identified in step S3 (i.e., the cracks in set D output in step S3), their multi-dimensional features are extracted to construct an explicit crack feature template library, providing a benchmark reference for hidden crack detection: For depth characterization extraction, the depth characterization corresponding to the visible cracks is extracted from the multi-task network in step S3. (256-dimensional, incorporating vibration-visual coupled semantics); For vibration and visual feature extraction, vibration features of visible cracks are extracted simultaneously. (256 dimensions, including kurtosis, energy percentage, etc.), visual features (512-dimensional, including information such as length and orientation angle); Infrared feature extraction was performed to acquire infrared thermal images of the visible crack region and obtain the temperature gradient. (Temperature gradient reflects the degree of spatial variation in the temperature field of the crack region), the average temperature difference between the crack and the background. This index quantifies the thermal difference between the cracked area and the normal background. It is the average temperature of the crack region. It is the average temperature of the background area; Organize the above features into a template library according to crack number. (m is the number of visible cracks).

[0079] For the candidate regions output in step S2 that were not identified as explicit cracks in step S3, the feature extraction process in step S3 is repeated to obtain the following multimodal features: For vibration features, vibration signals from candidate regions are extracted and processed using a 1D-CNN. ; For visual features, the image is cropped from the candidate region and then processed through a newly added convolutional layer in VGG16+ to obtain... ; Deep representation, will and Inputting the bidirectional attention and shared feature layer from step S3 yields a deep representation of the candidate region. ; By integrating the above data, a set of feature vectors for candidate regions is obtained. ); Then, infrared feature extraction is performed, and infrared thermal images of the candidate regions are acquired to obtain... and The process is as follows: Infrared thermograms were individually acquired for candidate regions not identified as visible cracks to obtain their temperature fields. Calculate the temperature gradient of the candidate region The calculation method is completely consistent with that for explicit crack regions, and the temperature gradient of candidate regions is extracted through differential extraction. , thus obtaining the gradient vector If cracks exist in the candidate region, its gradient vector will typically exhibit local anomalies (such as gradient abrupt changes).

[0080] Calculate the average temperature difference in candidate regions that were not identified as visible cracks. :

[0081] in, and These are the average temperature of the candidate crack region and the average background temperature, respectively. If this value deviates significantly from 0 (e.g., exceeds the threshold), the candidate region may be a real crack.

[0082] By quantifying feature differences and verifying physical laws, it is determined whether the candidate region meets the characteristics of a hidden crack (weak features but a pattern consistent with that of an explicit crack); By comparing the generated candidate region feature vector set with the explicit crack feature template library, feature differences are obtained. These feature differences include depth differences, vibration differences, and visual differences, as follows: For depth differences, calculate the depth representation vector of the candidate region. The deep representation template vector most similar to the template library Euclidean distance between :

[0083] Among them, the Euclidean distance reflects the straight-line distance between two vectors in space. The smaller the distance, the more similar the depth representation of the candidate region is to the depth representation of the explicit crack, and the more likely the candidate region is to be a hidden crack.

[0084] For vibration differences, calculate the vibration feature vector of the candidate region. Vibration feature template vector most similar to the template library The cosine similarity is calculated by subtracting 1 from the cosine similarity. :

[0085] in, It is a vibration difference. Cosine similarity measures the cosine value of the angle between two vectors. The smaller the angle, the higher the cosine similarity and the more matched the vibration characteristics. The smaller the value, the higher the vibration characteristic matching degree; For visual differences, similar to the calculation of vibration differences, the visual feature vectors of the candidate regions are calculated. The visual feature template vector most similar to the template library The cosine similarity is calculated, and 1 is subtracted from the cosine similarity to obtain the result. :

[0086] in, Visual differences The smaller the value, the higher the visual feature matching degree.

[0087] Candidate areas not identified as visible cracks are verified using infrared technology, and infrared thermal characteristic verification results are obtained (either one requirement is sufficient): Temperature gradient direction verification checks the temperature gradient direction of the candidate region against the temperature gradient template vector most similar to the template vector in the template library. Is the angle between the directions they represent less than or equal to 15°? If the angle meets the requirement, it indicates that the heat conduction direction of the crack is relatively consistent with the direction of the visible crack, supporting the judgment that the candidate area is a hidden crack.

[0088] Temperature difference verification is used to determine the average temperature difference of the candidate region. Is the temperature greater than or equal to 0.6°C? If so, it indicates that there is a change in thermal resistance in the candidate region due to the crack, providing evidence from a thermal perspective that the candidate region is a hidden crack.

[0089] A candidate area will be identified as a hidden crack only if all three of the following conditions are met: The first condition is that the depth difference meets the standard, which is the depth difference between the candidate region and the visible crack template. It must be greater than or equal to the difference in average depth between visible cracks. 25% of This indicates that although the candidate regions deviate somewhat from the characteristics of visible cracks in terms of depth representation, they still have similarities, which is consistent with the characteristic of hidden cracks having weak features but similar patterns.

[0090] The second condition is mode consistency and vibration characteristic differences. The difference in visual features must be less than or equal to 0.4. It must also be less than or equal to 0.4, that is... and This means that the vibration and visual patterns of the candidate region match the visible crack template well, eliminating completely irrelevant interference factors and further demonstrating its correlation with the crack.

[0091] It should be noted that the thresholds set for hidden crack identification—"depth difference ≥ 25% of the average depth difference of visible cracks, vibration difference ≤ 0.4, visual difference ≤ 0.4"—were determined based on a comprehensive analysis of large-sample experimental statistics, feature distribution analysis, and multi-scenario engineering verification. The depth difference threshold was based on the depth feature statistics of 120 visible cracks. Through comparative experiments of 80 hidden crack areas and 100 healthy areas, it was confirmed that this threshold can reduce the false positive rate of hidden cracks to 4.2% and cover 98% of real hidden cracks. The vibration difference threshold was derived from 150 visible cracks and 70 hidden cracks. Analysis of the cosine similarity distribution of the vibration characteristics of the crack shows that 0.4 is the optimal balance point between recognition sensitivity and specificity, corresponding to an recognition accuracy of 93%, which can effectively distinguish between healthy areas and cracked areas. The visual difference threshold was tested in 12 scenarios with different lighting, shooting angles, and crack widths. The setting of 0.4 can ensure that the visual feature similarity between hidden cracks and visible cracks is ≥85%, and the on-site verification accuracy in actual building inspection projects is stable at over 90%. At the same time, it can effectively eliminate interference from non-crack defects such as concrete spalling and water stains, ultimately achieving accurate identification of hidden cracks and avoiding misjudgment and omission.

[0092] The third condition is that the infrared verification passes, satisfying one of the conditions of the infrared verification. Candidate regions that meet the above criteria will be marked as hidden cracks. The following information will then be output: By obtaining the three-dimensional coordinates of hidden cracks through visual spatial mapping, the location of hidden cracks in the actual structure can be accurately located.

[0093] The damage index DI is calculated using the following formula:

[0094] in This represents the average temperature difference within the candidate regions. As the formula shows, a larger temperature difference corresponds to a higher damage index, reflecting the degree of damage from hidden cracks. Finally, all regions identified as hidden cracks are compiled to form a set of hidden cracks. .

[0095] This step supplements the identification of hidden cracks missed by visible cracks, and obtains information such as the location and degree of damage of hidden cracks. Together with the visible crack data in step S3, it constitutes a complete crack detection result, providing comprehensive crack data support for crack feature integration, grade determination and report generation in step S5.

[0096] Step S5: Based on the results of the identification of visible cracks and the results of the identification of hidden cracks, generate a structural crack detection report to realize the identification and location of structural cracks.

[0097] In real-world engineering scenarios, structures are often quite complex, and crack information obtained from different detection steps may differ in coordinate systems, feature description methods, and other aspects. To achieve unified and accurate crack analysis, the crack features are first standardized and integrated, as follows: Using the structural parameters determined in step S1 (such as structural dimensions and spatial location references), the three-dimensional coordinates of the visible cracks (set D) identified in step S3 and the hidden cracks (set H) detected in step S4 are uniformly transformed into the global coordinate system of the structure. In this way, all cracks have a consistent reference in spatial location, ensuring that the subsequent analysis and display of crack locations are accurate.

[0098] For visible cracks, in addition to the existing basic information, their damage index is further calculated:

[0099] in, It is the vibration amplitude, which reflects the impact of the vibration signal and is closely related to the changes in the structural mechanical properties caused by cracks; The crack length was obtained through visual feature extraction. The length of the largest crack in the sample set is used to normalize the crack length, making the damage index more comparable.

[0100] It should be noted that the weights for vibration amplitude (0.6) and crack length normalization (0.4) in the damage index formula were determined through comprehensive verification using experimental data fitting. Fifty samples of visible cracks with different vibration amplitudes and lengths were selected, and their structural stiffness degradation rate (a core mechanical indicator of damage degree) was tested. The vibration amplitude was analyzed using multiple linear regression. ) and normalized length ( The contribution of vibration amplitude to the stiffness degradation rate was fitted to a coefficient of 0.6 for the vibration amplitude and 0.4 for the length term. This weighting combination resulted in a damage index... The linear correlation coefficient with the actual stiffness degradation rate reached 0.92; After the above processing, a comprehensive crack set is generated. Each crack element in the set contains a wealth of information, such as three-dimensional coordinates (clearly indicating the crack's location within the structure), type (explicit or hidden, distinguishing the crack's manifestation), and damage index. (This measures the degree of damage caused by the cracks), laying the foundation for subsequent grade determination and visualization.

[0101] The crack level is determined as follows: To quickly determine the severity of cracks, they are classified into different levels based on the extent of damage: Visible cracks are determined based on the calculated damage index. To classify levels. When At a value of 0.3, it is classified as Level I, indicating that the damage caused by the cracks is relatively minor, and the structure can still maintain good performance in the short term; when 0.3... At 0.7, it is classified as Level II, indicating that the crack has caused moderate damage and its development needs to be closely monitored; when At a value of 0.7, it falls under Level III, indicating severe crack damage that could significantly impact structural safety, requiring immediate action.

[0102] Hidden cracks are classified as follows, based on the damage index DI and the significance of thermal anomalies estimated by infrared thermometry: When DI < 0.3, if the temperature difference ΔT between the crack and the background area is < 2°C, it is considered a low-damage, concealed crack with minor thermal anomalies. In this case, the damage caused by the crack is relatively minor, and the structure can maintain good performance in the short term, allowing for normal observation without special emergency treatment. If the temperature difference ΔT between the crack and the background area is ≥ 2°C, it is considered a low-damage, concealed crack with significant thermal anomalies. Although the damage is minor, the significant thermal anomalies indicate that the defect has a certain scale, requiring monitoring for further development and regular monitoring.

[0103] When 0.3 ≤ DI < 0.7, if the temperature difference ΔT from the background area is less than 2°C, it is considered a hidden crack with minor thermal anomaly and moderate damage. The crack has caused moderate damage, but the thermal anomaly is not obvious. Its development needs to be closely monitored, and some preventive maintenance measures can be taken. If the temperature difference ΔT from the background area is greater than or equal to 2°C, it is considered a hidden crack with significant thermal anomaly and moderate damage. The crack has caused moderate damage and the thermal anomaly is significant, with a more significant impact on the structure. The monitoring frequency should be increased, and the need for intervention measures should be assessed.

[0104] When DI ≥ 0.7, if the temperature difference ΔT between the crack and the background area is < 2°C, it is considered a hidden crack with minor thermal anomaly. The crack damage is severe, and even if the thermal anomaly is not obvious, it may still have a significant impact on structural safety, requiring timely analysis and appropriate measures. If the temperature difference ΔT between the crack and the background area is ≥ 2°C, it is considered a hidden crack with significant thermal anomaly. The crack damage is severe and the thermal anomaly is very significant, which is highly likely to cause more serious damage to the internal structure. Detailed testing and assessment should be carried out immediately, and reinforcement measures should be taken in a timely manner to ensure structural safety.

[0105] The identification and location results of cracks are clearly presented to relevant personnel through intuitive visualization and detailed reports. The process is as follows: Using the structure's BIM (Building Information Modeling), crack markers are overlaid on the model. Visible cracks are marked with solid red lines, and the line width increases with the Damage Index (DI), allowing for a visual distinction between visible cracks of varying damage levels. Hidden cracks are marked with dashed yellow lines, and the line width increases with crack depth, facilitating quick identification of the depth of hidden cracks. This method provides a highly intuitive display of crack identification and location within the structure, allowing relevant personnel to clearly understand the distribution and severity of cracks.

[0106] Output identification and location report: The crack statistics section provides a detailed count of the total number of cracks, as well as the proportion of visible and hidden cracks. It also lists the distribution of cracks of different grades, giving personnel a macroscopic understanding of the overall crack situation.

[0107] For the location details section, the three-dimensional coordinates of each crack are precisely listed. Combined with morphological feature information, the specific location and extension direction of the crack in the structure are clearly described, providing accurate location information for subsequent maintenance work.

[0108] This technical solution proposes a structural crack identification and localization method based on multimodal deep learning to address the challenge of accurately identifying both visible and hidden cracks using traditional detection methods. It achieves efficient and accurate crack detection through a fully structured design, combining prior building information to design a controllable micro-vibration excitation sequence, simultaneously acquiring vibration signals and image data to provide high-quality data for subsequent analysis. Anomaly correlation analysis between vibration signals and image data filters out highly suspicious candidate areas, narrowing the detection range. A multi-task convolutional neural network extracts vibration and visual dual-modal features, enhancing the accuracy of feature fusion through a bidirectional attention mechanism, while simultaneously classifying the existence and regressing the location coordinates of visible cracks. A visible crack feature template library is constructed, and high-suspicious hidden cracks are located by calculating feature differences. Infrared thermal features are introduced for cross-modal verification, enabling reliable identification of hidden cracks. Crack features are standardized and integrated to determine crack severity, and finally, annotated on the BIM model, outputting a detection report containing statistical and location information.

[0109] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, the phrase "comprising an element defined as..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0110] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their likenesses.

Claims

1. A method for structural crack identification and localization based on multimodal deep learning, characterized in that, Includes the following steps: Step S1: Collect prior information about the building structure, generate a controllable micro-vibration excitation sequence based on the prior information about the building structure, apply slight vibration to the building according to the controllable micro-vibration excitation sequence, and simultaneously collect the vibration signal and image data of the building; Step S2: Based on the vibration signal, mark the vibration anomaly area; based on the image data, mark the image anomaly area; by performing correlation analysis between the vibration anomaly area and the image anomaly area, select candidate crack areas; Step S3: Extract vibration signals and image data of the crack candidate region, construct a visible crack identification model based on a convolutional neural network, input the vibration signals and image data of the crack candidate region into the visible crack identification model, and output the visible crack identification result, which includes the visible crack region. Step S4: Based on the results of the visible crack identification, construct a visible crack feature template library, extract features from candidate regions that are not identified as visible cracks, generate a candidate region feature vector set, compare the candidate region feature vector set with the visible crack feature template library to obtain feature differences, verify the candidate regions that are not identified as visible cracks using infrared technology to obtain infrared thermal feature verification results, and combine the feature differences and infrared thermal feature verification results to obtain the hidden crack identification results. Step S5: Based on the identified explicit cracks and the identified hidden cracks, a structural crack detection report is generated to achieve structural crack identification and location. The construction of the explicit crack identification model includes: a vibration input branch extracts vibration feature vectors from vibration signals using a 1D-CNN; a visual input branch extracts visual feature vectors from image data using a 2D-CNN; a cross-correlation matrix of vibration and visual features is generated, where i and j are the indices of the vibration dimension and the visual dimension, respectively; vibration feature weight vectors and visual feature weight vectors are generated accordingly; the vibration and visual features are weighted separately and then concatenated to output fused features; the fused features are compressed and abstracted through a shared feature layer to obtain a deep representation; finally, the output layer, based on the deep representation, outputs the crack existence probability through a classification head; cracks with a probability greater than a threshold are identified as explicit cracks; and the boundary coordinates of the explicit cracks are output through a regression head.

2. The structural crack identification and localization method based on multimodal deep learning according to claim 1, characterized in that, Collect prior information about the building structure, and based on this prior information, generate a controllable micro-vibration excitation sequence, including: The prior information of the target building structure was collected as follows: First, collect the design documents, as-built reports, and operation and maintenance records of the target building and organize them into a basic information database; The controllable micro-vibration excitation sequence based on structural partitioning is designed as follows: First, the excitation parameter constraints are determined for each sub-region. Based on its mechanical parameters Calculate the constraint range of the excitation parameters: The frequency constraint is estimated using the formula for the natural frequency of the component: ; in The span of the component is expressed in meters (m). This refers to the material density, in units of... ; The cross-sectional area is expressed in meters (m²). 2 ; I is the elastic modulus of the material, measured in Pa; I is the moment of inertia of the cross section, measured in meters. 4 Determine the excitation frequency range , usually take ; Amplitude constraint determines the maximum excitation force through preliminary experiments. Ensure structural strain ; Duration constraint is the duration of single-region excitation. ; Based on the above-mentioned excitation parameter constraints, for each sub-region Generate a dedicated incentive sequence .

3. The structural crack identification and localization method based on multimodal deep learning according to claim 2, characterized in that, According to the controllable micro-vibration excitation sequence, slight vibrations are applied to the building, and vibration signals and image data of the building are collected simultaneously, including: According to the excitation sequence Vibrations were applied sequentially to each sub-region, and the data was recorded synchronously as follows: Vibration sensors collect acceleration signals Stored sequentially in .csv format; Camera captures continuous image sequences Store as a .png file, along with the associated sub-region ID and excitation parameters.

4. The structural crack identification and localization method based on multimodal deep learning according to claim 3, characterized in that, Based on the vibration signal, abnormal vibration areas are marked; based on the image data, abnormal image areas are marked, including: The correlation analysis algorithm is used to identify regions that are in phase with the anchor point vibration but have abnormal intensity. The process of marking the vibration anomaly region is as follows: For each monitoring point, the preprocessed vibration acceleration signal is extracted. sequence Simultaneously, extract the preprocessed vibration acceleration signal sequence of the anchor point. Calculate the Pearson correlation coefficient between the two. : ; in, It is the preprocessed vibration acceleration signal from the monitoring point. In the Each sampling time The value of , The vibration acceleration signal of the anchor point after preprocessing is at the 1st Each sampling time The value of , This represents the number of signal sampling points. It is the average value of the signals from the monitoring points. It is the mean of the anchor signal; the numerator is the covariance of the two signals, and the denominator is the product of the standard deviations of the two signals; like The vibration in this area was determined to be in phase with that of the anchor point, and it was marked as an abnormal vibration area. Its spatial coordinate range was recorded. ; Based on preprocessed normalized image sequences The process of marking abnormal regions in an image using threshold analysis is as follows: For each frame of the image, contours are extracted using Canny edge detection, retaining continuous edges with a length > 10 pixels. For a given edge detected in the k-th frame... Suppose it contains Pixels, coordinates are Then the center coordinates of the edge are : ; in, It is the first coordinates of pixels in These are horizontal pixel coordinates. These are vertical pixel coordinates; It is the first frame edges The coordinates of the horizontal center point, through the... The average of the horizontal coordinates of each pixel is obtained; It is the first frame edges The coordinates of the vertical center point, through the... The average of the vertical coordinates of each pixel is used to obtain the value, where m is the number of pixels. For the same edge matched in two consecutive frames, calculate the coordinates of the center point. and The Euclidean distance between the two is: ; in, It is the Euclidean distance between the center points of the edges of two frames. These are the coordinates of the center point of the edge in the (k+1)th frame; The timing stability is determined as follows: For 20 consecutive frames of images, calculate the displacement between every two adjacent frames. Take the average displacement. k is the frame index; like Pixels are identified as temporally stable edges and marked as image anomaly regions. The pixel coordinates of these anomaly regions are then converted to 3D architectural coordinates via spatial mapping and recorded. .

5. The structural crack identification and localization method based on multimodal deep learning according to claim 4, characterized in that, Vibration signals and image data of candidate crack regions are extracted. A visible crack identification model is constructed based on a convolutional neural network. The vibration signals and image data of the candidate crack regions are input into the visible crack identification model, and the visible crack identification results are output, including: Based on convolutional neural networks, a visible crack identification model is constructed as follows: The input layer is divided into two branches; One branch is the vibration input. For the crack candidate region output in step S2, the acceleration signal of the corresponding vibration sensor is extracted, and high-frequency noise is removed by a 50Hz low-pass filter to obtain the preprocessed signal. ; Another branch uses visual input, mapping the 3D coordinates of the crack candidate region to the image plane and cropping the ROI region. Enhance crack edges through adaptive threshold segmentation; The feature extraction branch is divided into a vibration feature extraction branch and a visual feature extraction branch. The bidirectional attention fusion layer is as follows: First, the correlation matrix is ​​calculated to generate the cross-correlation matrix of vibration-visual features. : ; Where i and j are the indices of the vibration dimension and the visual dimension, respectively, quantifying the correlation between the i-th dimension of vibration and the j-th dimension of vision; Weighted fusion is performed, with vibration and visual features weighted separately and then stitched together to output a 768-dimensional fused feature. : ; in, It is the weight vector of vibration characteristics. The weight vector of vibration characteristics, Element-wise multiplication For dimensional splicing, It is the vibration characteristic vector. It is a visual feature vector; By using a shared feature layer, the features output from the vibration feature extraction branch and the visual feature extraction branch are further fused, compressed, and abstracted to obtain more representative deep features. ; The final output layer is a multi-task output head, which is based on the deep representation output by the shared feature layer. It simultaneously performs two tasks: determining the existence of cracks and locating cracks, in order to meet the actual needs of crack detection. It is divided into a classification head and a regression head, and the classification head is output first, followed by the regression head. The classification header output process is as follows: This part uses a fully connected layer to represent the 256-dimensional depth. The output is converted to one dimension, and then the Sigmoid activation function is used to map the output value to the interval [0, 1] to obtain the probability of crack existence. , The range of values ​​for is [0, 1]. The closer the probability is to 1, the more likely the input region is to be a visible crack; the model takes the output probability. The candidate cracks were determined to be explicit cracks; The regression head output process is as follows: Using a fully connected layer, the classification head determines the 256-dimensional depth characterization of explicit cracks. Convert to 4D output, outputting the coordinates of the explicit crack boundary. .

6. The structural crack identification and localization method based on multimodal deep learning according to claim 5, characterized in that, Based on the results of explicit crack identification, a explicit crack feature template library is constructed, including: Based on the visible cracks identified in step S3, their multi-dimensional features are extracted to construct a visible crack feature template library, providing a benchmark reference for the detection of hidden cracks. For depth characterization extraction, the depth characterization corresponding to the visible cracks is extracted from the multi-task network in step S3. ; For vibration and visual feature extraction, vibration features of visible cracks are extracted simultaneously. Visual features ; Infrared feature extraction was performed to acquire infrared thermal images of the visible crack region and obtain the temperature gradient. The temperature gradient reflects the degree of spatial variation in the temperature field within the crack region and the average temperature difference between the crack and the background. , Quantify the thermal difference between the crack region and the normal background. It is the average temperature of the crack region. It is the average temperature of the background area; Organize the above features into a template library according to crack number. .

7. The structural crack identification and localization method based on multimodal deep learning according to claim 6, characterized in that, Feature extraction is performed on candidate regions not identified as overt cracks to generate a candidate region feature vector set. The candidate region feature vector set is then compared with an overt crack feature template library to obtain feature differences, including: For the candidate regions output in step S2 that were not identified as explicit cracks in step S3, the feature extraction process in step S3 is repeated to obtain the following multimodal features: For vibration features, vibration signals from candidate regions are extracted and processed using a 1D-CNN. ; For visual features, the image is cropped from the candidate region and then processed through a newly added convolutional layer in VGG16+ to obtain... ; Deep representation, will and Inputting the bidirectional attention and shared feature layer from step S3 yields a deep representation of the candidate region. ; By integrating the above data, a set of feature vectors for candidate regions is obtained. ); By comparing the generated candidate region feature vector set with the explicit crack feature template library, feature differences are obtained. These feature differences include depth differences, vibration differences, and visual differences, as follows: For depth differences, calculate the depth representation vector of the candidate region. The deep representation template vector most similar to the template library Euclidean distance between : ; Among them, the Euclidean distance reflects the straight-line distance between two vectors in space. The smaller the distance, the more similar the depth representation of the candidate region is to the depth representation of the explicit crack, and the more likely the candidate region is to be a hidden crack. For vibration differences, calculate the vibration feature vector of the candidate region. Vibration feature template vector most similar to the template library The cosine similarity is calculated by subtracting 1 from the cosine similarity. : ; in, It refers to vibration differences. Cosine similarity measures the cosine of the angle between two vectors. The smaller the angle, the higher the cosine similarity, and the more matched the vibration characteristics. The smaller the value, the higher the vibration characteristic matching degree; For visual differences, similar to the calculation of vibration differences, the visual feature vectors of the candidate regions are calculated. The visual feature template vector most similar to the template library The cosine similarity is calculated, and 1 is subtracted from the cosine similarity to obtain the result. : ; in, Visual differences The smaller the value, the higher the visual feature matching degree.

8. The structural crack identification and localization method based on multimodal deep learning according to claim 7, characterized in that, Candidate regions not identified as visible cracks are verified using infrared technology, yielding infrared thermal characteristic verification results, including: Candidate regions not identified as visible cracks were verified using infrared technology, and the infrared thermal feature verification results were obtained: Temperature gradient direction verification checks the temperature gradient direction of the candidate region against the temperature gradient template vector most similar to the template vector in the template library. If the angle between the directions represented is less than or equal to 15°, it means that the heat conduction direction of the crack is relatively consistent with the direction of the visible crack, which supports the judgment that the candidate area is a hidden crack. Temperature difference verification is used to determine the average temperature difference of the candidate region. If the temperature is greater than or equal to 0.6°C, it indicates that there is a change in thermal resistance in the candidate region due to the crack, providing evidence from a thermal perspective that the candidate region is a hidden crack.

9. A structural crack identification and localization method based on multimodal deep learning according to claim 8, characterized in that, Combining the aforementioned feature differences and infrared thermal feature verification results, the hidden crack identification results are obtained, including: A candidate area will be identified as a hidden crack only if all three of the following conditions are met: The first condition is that the depth difference meets the standard, which is the depth difference between the candidate region and the visible crack template. It must be greater than or equal to the difference in average depth between visible cracks. 25% of ; The second condition is mode consistency and vibration characteristic differences. The difference in visual features must be less than or equal to 0.

4. It must also be less than or equal to 0.4, that is... and ; The third condition is that the infrared verification is passed, which is one of the conditions for infrared verification. Infrared verification is divided into temperature gradient direction verification and temperature difference verification. Temperature gradient direction verification checks the temperature gradient direction of the candidate region against the temperature gradient template vector most similar to the template vector in the template library. If the angle between the directions represented is less than or equal to 15°, it means that the heat conduction direction of the crack is relatively consistent with the direction of the visible crack, which supports the judgment that the candidate area is a hidden crack. Temperature difference verification is used to determine the average temperature difference of the candidate region. If the temperature is greater than or equal to 0.6°C, it indicates that there is a change in thermal resistance in the candidate region due to the crack, providing evidence from a thermal perspective that the candidate region is a hidden crack. Candidate areas that meet the above criteria will be marked as hidden cracks.

10. A structural crack identification and localization method based on multimodal deep learning according to claim 9, characterized in that, Based on the identified visible and hidden cracks, a structural crack detection report is generated, including: The results of crack identification and location are presented to relevant personnel through intuitive visualization and detailed reports. The process is as follows: Using structural BIM, crack markers are overlaid on the model. Visible cracks are marked with solid red lines, and the line width increases with the damage index. The width of the line increases with the increase of the damage index DI, which makes it easy to visually distinguish visible cracks of different damage levels; for hidden cracks, a yellow dashed line is used to mark them, and the line width increases with the increase of the damage index DI, which makes it easy to quickly identify the condition of hidden cracks. The identification and location report is as follows: For the crack statistics section, a detailed count of the total number of cracks, as well as the proportion of visible cracks and hidden cracks, is provided. The distribution of cracks of different grades is also listed to give personnel a macroscopic understanding of the overall crack situation. For the location details section, the three-dimensional coordinates of each crack are precisely listed. Combined with morphological feature information, the specific location and extension direction of the crack in the structure are clearly described, providing accurate location information for subsequent maintenance work.