Building crack scanning monitoring method based on deep learning
By constructing a multi-level spatial grid structure and a multi-domain input field, combined with an improved ANO model and a cross-scale attention mechanism, the problems of missed detection and misjudgment in the scanning of cracks on building facades were solved, realizing continuous and systematic scanning and high-precision crack identification of building facades.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- EAST CHINA JIAOTONG UNIVERSITY
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies are insufficient for continuous and systematic crack scanning and identification on building facades, especially in large buildings where there is a risk of missed detection and a high rate of false positives. Furthermore, traditional methods lack the ability to acquire information about structural vibration behavior, making it difficult to accurately identify early or hidden cracks.
A multi-level spatial grid structure is constructed, and a multi-domain input field is introduced to encode the time-frequency characteristics of acoustic vibration response and spatial grid characteristics in a unified manner. An improved ANO model is used to realize the synchronous modeling of the coarse grid structural response and the local characteristics of cracks in the fine grid. A crack risk field is generated based on a cross-scale attention mechanism, and the structural response information of acoustic vibration excitation and laser vibration radar is utilized.
It enables continuous scanning and identification of building facades, improving identification accuracy and scanning range. It can finely detect crack risks on complex facades, output crack risk heat maps and area contours, and significantly improve the accuracy and stability of scanning monitoring.
Smart Images

Figure CN122017018A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of building facade inspection, and more particularly to a deep learning-based method for scanning and monitoring building cracks. Background Technology
[0002] Crack detection on building facades typically relies on manual inspections, close-up photography, or crack identification methods based on two-dimensional images. Manual inspections are limited by the experience of the inspectors and struggle to continuously and systematically scan and identify the overall condition of large facades, leading to risks of missed detections and subjectivity. Traditional image recognition methods rely solely on surface texture information, failing to reflect the structural response characteristics of cracks. They are also sensitive to changes in lighting, surface contamination, and material differences, resulting in a high false positive rate. Furthermore, detection methods based on a single image modality lack the ability to capture structural vibration behavior, making it difficult to accurately identify early-stage or hidden cracks.
[0003] While existing acoustic and vibration detection technologies can be used for structural defect identification, they typically rely on discrete sampling point measurements, resulting in fragmented acquisition methods and a lack of continuous scanning capability across the entire building facade. Furthermore, traditional vibration signal analysis methods often remain at the level of feature extraction in the time, frequency, or time-frequency domains, making it difficult to deeply integrate multi-source response features with spatial distribution information and establish a unified characterization of structural responses and local anomalies at different scales. In addition, existing multi-scale analysis methods usually employ fixed resolution or single-scale network structures, failing to adequately utilize the correlation between coarse-scale structural changes and fine-scale local cracks in complex facades.
[0004] In the fusion processing of multi-source vibration response data and spatial image coordinates, existing technologies generally lack a holistic modeling method for the structural state of various areas of a building facade, especially lacking a cross-scale analysis mechanism that can simultaneously infer crack risk based on coarse and fine-scale structural responses. Existing deep learning models, when handling the task of fusing facade vibration response with spatial features, often suffer from insufficient sensitivity to physical parameters, weak spatial structural coupling, and inaccurate identification of local high-risk areas, thus limiting the application of scanning crack monitoring technology in practical engineering.
[0005] Therefore, how to provide a deep learning-based method for monitoring building cracks is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] One objective of this invention is to propose a deep learning-based method for scanning and monitoring building cracks. This invention constructs a multi-level spatial grid structure, introduces a multi-domain input field, and uniformly encodes the time-frequency characteristics of acoustic and vibration responses, spatial grid features, and scanning physical parameters. An improved ANO model is used to simultaneously model the coarse-grid structural response and the local crack features of the fine-grid. Furthermore, a crack risk field is generated based on a cross-scale attention mechanism, thereby enabling continuous scanning and identification of building facades. This invention fully utilizes the structural response information from acoustic and vibration excitation and laser vibration radar, and introduces innovative structures such as physical modulation, multi-domain operators, and cross-scale fusion. This allows for refined detection of crack risks on complex facades, offering advantages such as high identification accuracy, large scanning range, and strong physical consistency of the model.
[0007] A deep learning-based method for monitoring building cracks according to an embodiment of the present invention includes the following steps:
[0008] A swept-frequency acoustic vibration excitation signal is radiated onto the exterior of the building, and the acoustic vibration response time-series data of the exterior of the building are collected to construct an acoustic vibration response time-series data set;
[0009] Spectral analysis is performed on the acoustic vibration response time-series data set, and a time-frequency feature tensor is constructed based on the spectral analysis results to generate a time-frequency feature set;
[0010] Based on the measurement point location records of the acoustic vibration response time series data set, a multi-level spatial grid structure is constructed, each measurement point is assigned to a grid cell in the multi-level spatial grid structure, the vibration energy characteristics and statistical characteristics of each grid cell are calculated, and a spatial grid feature set is generated.
[0011] The time-frequency feature set is mapped to the time-frequency feature domain input field, and the spatial grid feature set is mapped to the spatial heat map domain input field, which are then combined into a multi-domain input field set.
[0012] An improved ANO model is constructed, using a multi-level spatial grid structure as the operator's domain, performing multi-domain attention operator processing, defining coarse grid operators and fine grid operators, and generating a coarse grid structure response feature field and a fine grid crack local feature field.
[0013] A cross-scale attention mechanism is established based on a multi-level spatial grid structure. The response feature field of the coarse grid structure is used to modulate the local feature field of the crack in the fine grid to generate a crack risk field.
[0014] Based on the crack risk field, threshold judgment and connectivity analysis are performed on the crack risk value corresponding to each grid cell in the crack risk field to generate the marking results of crack areas on the building facade.
[0015] Optionally, the construction of the acoustic-vibration response time-series data set includes:
[0016] Set up a subset of sound source parameters, a subset of radar parameters, and a subset of scanning control parameters. Perform numerical range setting, step resolution setting, and unit unification processing on the three subsets respectively to generate a set of scanning parameters.
[0017] Based on the geometric contour information and scanning mode of the building facade, the scanning angle and scanning speed are obtained from the subset of scanning control parameters. A preset scanning path covering the detection area is generated in the coordinate system of the building facade. The scanning trajectory point sequence and corresponding time sequence of the laser vibration radar are determined to form the preset scanning path data.
[0018] Based on the set of scanning parameters and preset scanning path data, the directional sound source is controlled to radiate frequency-sweeping acoustic and vibration excitation signals to the exterior of the building, and the laser vibration radar is controlled to collect acoustic and vibration response time-series data corresponding to the time series in sequence, generating the original acoustic and vibration response time-series data.
[0019] The original acoustic and vibration response time series data were synchronized and the measurement points were numbered to construct an acoustic and vibration response time series data set.
[0020] Optionally, the generation of the time-frequency feature set includes:
[0021] Perform spectral analysis on the acoustic and vibration response time-series data set, including Fourier transform, time-frequency transform, noise estimation, validity assessment, and power spectrum calculation;
[0022] The spectral analysis results are reorganized and stacked according to the measurement point dimension, time dimension, and frequency dimension to construct a time-frequency feature tensor;
[0023] The data segments corresponding to each measurement point in the time-frequency feature tensor are extracted into measurement point-level feature units and organized to form a time-frequency feature set.
[0024] Optionally, the generation of the spatial grid feature set includes:
[0025] Based on the measurement point location records in the acoustic vibration response time series data set, the polar coordinate measurement point coordinates output by the laser vibration radar are paired one by one with the rectangular coordinates of the building facade image. The polar coordinate measurement point coordinates are then transformed to obtain the spatial coordinates of the measurement points, and a set of spatial coordinates of the measurement points is generated.
[0026] Under the building facade coordinate system corresponding to the spatial coordinate set of measurement points, a coarse grid and a fine grid are divided according to a preset grid size. Coarse grid index and fine grid index are set respectively to construct a multi-level spatial grid structure covering the building facade.
[0027] The time-frequency features corresponding to each measuring point in the time-frequency feature set are assigned to the corresponding coarse and fine grid cells according to the position of the measuring point's spatial coordinates in the multi-level spatial grid structure. Based on the time-frequency features corresponding to the measuring points assigned to the grid cells, the vibration energy features and statistical features of the grid cells are calculated, and a spatial grid feature set is generated.
[0028] Optionally, the generation of the multi-domain input field set includes:
[0029] Based on the time-frequency feature set, the spatial coordinate set of measurement points, and the multi-level spatial grid structure, for each coarse and fine grid cell in the multi-level spatial grid structure, the measurement points belonging to the grid cell are determined. The time-frequency features of the belonging measurement points in the time-frequency feature set are weighted and normalized according to the preset aggregation rules. The processing results are written into the corresponding grid cell. After all grid cells are written, the time-frequency feature domain input field is formed.
[0030] In a multi-level spatial grid structure, the vibration energy characteristics and statistical characteristics of the grid cell are read from the spatial grid feature set, encoded into the spatial feature vector of the grid cell, and written into the corresponding grid cell. After all grid cells are written, the spatial heat map domain input field is formed.
[0031] Numerical normalization and encoding are performed on the set of scanning parameters. The encoded scanning parameters are arranged into a scanning parameter vector. In the grid cells of the multi-level spatial grid structure, the scanning parameter vector is copied into the operator kernel physical modulation vector corresponding to the grid cell and arranged to generate a scanning parameter modulation field.
[0032] Based on the time-frequency feature domain input field, the spatial heat map domain input field, and the scanning parameter modulation field, the three input fields are spliced together and organized into a multi-domain input field set.
[0033] Optionally, the construction and use of the improved ANO model includes:
[0034] Based on a multi-domain input field set and a multi-level spatial grid structure, input channels are set for the multi-domain input field set on grid nodes aligned with the multi-level spatial grid structure. The three input channels are mapped and aligned on the multi-level spatial grid structure to construct the multi-domain operator input structure of the improved ANO model. A multi-domain attention operator structure is set, with the multi-level spatial grid structure as the operator scope.
[0035] In the improved ANO model, a multi-domain operator input structure is used as input. A time-frequency operator branch is constructed on the input channel corresponding to the time-frequency feature domain input field, and a spatial operator branch is constructed on the input channel corresponding to the spatial heatmap domain input field. Attention weight calculation units are set up, and the attention weights of each grid node in the time-frequency operator branch and the spatial operator branch are jointly modulated based on the operator kernel physical modulation vector.
[0036] On the coarse and fine grids of the multi-level spatial grid structure, coarse and fine grid operators are configured based on the coarse and fine grid indices, respectively, for the modulated time-frequency operator branch and spatial operator branch;
[0037] Perform coarse mesh operator operations to generate a coarse mesh structural response feature field; perform fine mesh operator operations to generate a fine mesh crack local feature field.
[0038] Optionally, the generation of the crack risk field includes:
[0039] In the coarse mesh structure response feature field, a coarse mesh index is assigned to each coarse mesh node, and in the fine mesh crack local feature field, a fine mesh index is assigned to each fine mesh node. Based on the spatial inclusion relationship between the coarse mesh index and the fine mesh index, the coarse mesh node to which each fine mesh node belongs is determined, and a set of coarse and fine mesh mapping relationships is generated.
[0040] Based on the set of coarse and fine mesh mapping relationships, the coarse mesh structural response characteristics of each coarse mesh node and the local characteristics of fine mesh cracks of the corresponding fine mesh node are read. Cross-scale correlation measures are calculated, normalization is performed, and a cross-scale attention weight set is generated.
[0041] A cross-scale attention mechanism is constructed on a multi-level spatial grid structure. At each fine grid node, the cross-scale attention weight is used to perform a weighted combination of the coarse grid structure response features of the corresponding coarse grid node and the local features of the fine grid cracks of the fine grid node, thereby generating crack risk values for the fine grid nodes and arranging them to generate a crack risk field.
[0042] Optionally, the generation of the building facade crack area marking results includes:
[0043] Based on the distribution of the crack risk field on a multi-level spatial grid structure, the crack risk value of each grid cell is read from the crack risk field, a threshold judgment is performed, crack candidate grid cells and non-crack grid cells are marked, the marking results of crack candidate grid cells and non-crack grid cells are organized, and a crack candidate marking matrix is generated.
[0044] Based on the crack candidate marker matrix and multi-level spatial grid structure, the spatial adjacency relationship between grid cells is determined, connectivity analysis is performed on the crack candidate grid cells, the set of crack connected regions composed of spatially adjacent crack candidate grid cells is identified, the number of grid cells and the cumulative sum of crack risk values in the crack connected regions are calculated, and sets with the number of grid cells and the cumulative sum of crack risk values below the threshold are removed to generate a set of effective crack connected regions.
[0045] Based on the set of effective crack connected regions, the mesh cells belonging to the same effective crack connected region are aggregated into a single crack region. The outer contour extraction operation is performed on the set of fine mesh cells in each crack region to generate the corresponding crack region contour.
[0046] For each crack region, calculate the crack area estimation result, generate a crack area estimation result set, and store the crack region outline and crack area estimation result set corresponding to the effective crack connected region set to generate the building facade crack region marking result;
[0047] The crack risk values in the crack risk field are color-coded to generate a crack risk heat map of the building facade. The crack area outline and crack area estimation results are overlaid and displayed, and a building crack scanning monitoring report containing the crack risk heat map, crack area outline and crack area estimation results is output.
[0048] The beneficial effects of this invention are:
[0049] First, by constructing a multi-level spatial grid structure, this invention maps the time-frequency characteristics of acoustic and vibration responses to the spatial distribution of building facades in a unified manner, thereby achieving continuous and systematic scanning of a large area of the facade. This effectively overcomes the limitations of traditional detection methods that rely on discrete measurement points and are difficult to obtain the overall response.
[0050] Secondly, this invention introduces a multi-domain input field, which integrates time-frequency features, spatial grid features, and scanning physical parameters into the improved ANO model. Through physical modulation and multi-domain attention mechanisms, the model's sensitivity to the differences in vibration behavior in different structural regions is enhanced, enabling the structural-level response and local crack features to be accurately modeled simultaneously.
[0051] Furthermore, this invention constructs a cross-scale attention mechanism between the coarse-scale structural response feature field and the fine-scale crack local feature field, which can fully utilize the guiding role of the coarse-scale structural state on the fine-scale crack risk, achieve high-precision generation of the crack risk field, and make crack region identification more reliable.
[0052] Furthermore, by using threshold judgment and connectivity analysis of the crack risk field, this invention can finely mark crack regions and output risk heat maps and crack outlines, thereby significantly improving the accuracy, stability, and engineering applicability of crack scanning monitoring. Attached Figure Description
[0053] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0054] Figure 1 This is an overall flowchart of a deep learning-based building crack scanning and monitoring method proposed in this invention;
[0055] Figure 2 This is a schematic diagram illustrating the construction of the multi-domain input field in this invention;
[0056] Figure 3 This is a schematic diagram of the cross-scale attention mechanism of the improved ANO model in this invention;
[0057] Figure 4 , Figure 5 These are comparison images from the first detection task in the facade crack scanning and monitoring mission of a building with an exterior wall material of mortar and paint, as shown in the example.
[0058] Figure 6 , Figure 7 These are comparison images from the second detection task in the facade crack scanning and monitoring task of a building with an exterior wall material of mortar and paint, as shown in the example.
[0059] Figure 8 , Figure 9 , Figure 10 These are comparison images from the first detection task in the facade crack scanning and monitoring mission of a building with ceramic tile exterior walls, as shown in the example.
[0060] Figure 11 , Figure 12 These are comparison images from the second detection task in the facade crack scanning and monitoring task of a building with ceramic tile exterior wall material, as shown in the example. Detailed Implementation
[0061] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0062] refer to Figure 1-3 A deep learning-based method for detecting building cracks includes the following steps:
[0063] Set a set of scanning parameters, including sound source power, sweep frequency excitation signal, playback time, radar focal length, scanning mode, scanning angle, scanning speed, frequency range, and amplitude range. Control the directional sound source to radiate the sweep frequency acoustic vibration excitation signal to the building facade. Control the laser vibration radar to collect the acoustic vibration response time series data of the building facade on the preset scanning path and construct the acoustic vibration response time series data set.
[0064] Spectral analysis processing is performed on the acoustic vibration response time series data set, including Fourier transform, time-frequency transform, noise estimation, validity judgment and power spectrum calculation. Based on the scanning parameter set and the spectral analysis results, a time-frequency feature tensor is constructed to generate a time-frequency feature set.
[0065] Based on the measurement point location records of the acoustic vibration response time series data set, the polar coordinate measurement point coordinates output by the laser vibration radar are transformed with the rectangular coordinates of the building facade image to obtain a set of measurement point spatial coordinates that correspond one-to-one with each measurement point in the time-frequency feature set. At the same time, under the unified building facade coordinate system, coarse and fine grids are divided according to the preset grid size to construct a multi-level spatial grid structure containing coarse and fine grid indices. Based on the measurement point spatial coordinate set and the time-frequency feature set, each measurement point is assigned to a grid cell in the multi-level spatial grid structure. The vibration energy characteristics and statistical characteristics of each grid cell are calculated to generate a spatial grid feature set.
[0066] Based on the set of scanning parameters, the set of time-frequency features, the set of spatial coordinates of measurement points, the multi-level spatial grid structure, and the set of spatial grid features, the set of time-frequency features is mapped into a time-frequency feature domain input field aligned with the coarse and fine grids according to the multi-level spatial grid structure. The set of spatial grid features is mapped into a spatial heat map domain input field aligned with the multi-level spatial grid structure. The set of scanning parameters is encoded into a physical modulation vector field of the operator kernel to form a scanning parameter modulation field. The time-frequency feature domain input field, the spatial heat map domain input field, and the scanning parameter modulation field are combined into a multi-domain input field set, and the multi-domain input field set is input into the improved ANO model.
[0067] An improved ANO model is constructed, using a multi-level spatial grid structure as the operator domain. Multi-domain attention operator processing is performed on the multi-domain input field set. Time-frequency operator branches and spatial operator branches are constructed in the time-frequency feature domain input field and the spatial heatmap domain input field, respectively. The attention weights of the time-frequency operator branches and spatial operator branches are jointly modulated by the modulation vector generated by the scanning parameter modulation field. Coarse grid operators and fine grid operators are defined on the coarse grid and fine grid of the multi-level spatial grid structure, respectively. The coarse grid operator is applied on the coarse grid to generate the coarse grid structure response feature field, and the fine grid operator is applied on the fine grid to generate the fine grid crack local feature field.
[0068] A cross-scale attention mechanism is established between the coarse mesh structural response feature field and the fine mesh crack local feature field based on the mesh index of the multi-level spatial mesh structure. The coarse mesh structural response feature field is used to modulate the fine mesh crack local feature field through the cross-scale attention mechanism to generate a crack risk field defined on the multi-level spatial mesh structure.
[0069] Based on the distribution of the crack risk field on a multi-level spatial grid structure, threshold judgment and connectivity analysis are performed on the crack risk value corresponding to each grid cell in the crack risk field to generate the crack area marking results of the building facade and output a building crack scanning and monitoring report containing crack risk heat map, crack area outline and crack area estimation results.
[0070] In this embodiment, the construction of the acoustic vibration response time-series data set includes:
[0071] Set up a subset of sound source parameters, a subset of radar parameters, and a subset of scanning control parameters. The subset of sound source parameters includes sound source power, frequency sweep excitation signal, and playback time. The subset of radar parameters includes radar focal length. The subset of scanning control parameters includes scanning mode, scanning angle, scanning speed, frequency range, and amplitude range. Perform numerical range setting, step resolution setting, and unit unification processing on the subset of sound source parameters, the subset of radar parameters, and the subset of scanning control parameters respectively to generate a set of scanning parameters.
[0072] Based on the geometric contour information and scanning mode of the building facade, the scanning angle and scanning speed are obtained from the subset of scanning control parameters. A preset scanning path covering the detection area is generated in the coordinate system of the building facade. The scanning trajectory point sequence and corresponding time sequence of the laser vibration radar are determined on the preset scanning path to form the preset scanning path data.
[0073] Based on the set of scanning parameters and preset scanning path data, the directional sound source is controlled to radiate the frequency sweeping sound and vibration excitation signal to the building facade according to the sound source power, frequency sweeping excitation signal and playback time. The laser vibration radar is controlled to collect the sound and vibration response time sequence data corresponding to the time sequence on the building facade according to the radar focal length and the scan trajectory point sequence in the preset scanning path data, and generate the original sound and vibration response time sequence data.
[0074] The original acoustic and vibration response time series data are time-synchronized and organized according to the time series and the scan trajectory point series, and the measurement point numbers are organized. Abnormal data segments that do not match the preset scan path data are removed. The remaining original acoustic and vibration response time series data are archived and stored according to the coordinate system of the building facade to construct an acoustic and vibration response time series data set.
[0075] In this embodiment, the generation of the time-frequency feature set includes:
[0076] Perform spectral analysis on the acoustic and vibration response time-series data set, including Fourier transform, time-frequency transform, noise estimation, validity assessment, and power spectrum calculation;
[0077] The spectral analysis results are reorganized and stacked according to the measurement point dimension, time dimension, and frequency dimension to construct a time-frequency feature tensor expanded in the measurement point dimension, time dimension, and frequency dimension.
[0078] The data segments corresponding to each measurement point in the time-frequency feature tensor are extracted into measurement point-level feature units and organized to form a time-frequency feature set.
[0079] In this embodiment, the generation of the spatial grid feature set includes:
[0080] Based on the measurement point location records in the acoustic vibration response time series data set, the polar coordinate measurement point coordinates output by the laser vibration radar are paired one by one with the rectangular coordinates of the building facade image. The polar coordinate measurement point coordinates are transformed under a unified building facade coordinate system to obtain the measurement point spatial coordinates that correspond one-to-one with each measurement point in the time-frequency feature set. The coordinate transformation results are organized according to the measurement point identifier to generate a measurement point spatial coordinate set.
[0081] Under the coordinate system of the building facade corresponding to the spatial coordinate set of the measurement points, a coarse grid and a fine grid are divided according to the preset grid size. Coarse grid index and fine grid index are set in the coarse grid and the fine grid respectively. Based on the coarse grid index and the fine grid index, a multi-level spatial grid structure covering the building facade is constructed. The multi-level spatial grid structure is used as the spatial division basis for subsequent grid unit division and feature calculation.
[0082] Based on the set of spatial coordinates of measurement points and the set of time-frequency features, the time-frequency features corresponding to each measurement point in the time-frequency feature set are assigned to the corresponding coarse and fine grid cells according to the position of the measurement point's spatial coordinates in the multi-level spatial grid structure. Within each grid cell, the vibration energy features and statistical features of the grid cell are calculated based on the time-frequency features corresponding to the measurement points assigned to the grid cell. The statistical features include at least the mean amplitude, amplitude variance, and amplitude peak value. The vibration energy features and statistical features of each grid cell are organized according to the coarse grid index and the fine grid index to generate a spatial grid feature set.
[0083] The calculation of vibration energy characteristics and statistical characteristics includes: extracting the amplitude value from the time-frequency characteristics corresponding to the measurement point belonging to the grid cell within each grid cell; performing weighted summation on the amplitude values in the time and frequency dimensions to obtain the vibration energy characteristics representing the overall vibration intensity of the grid cell; calculating the arithmetic mean of the set of amplitude values in the time and frequency dimensions to obtain the amplitude mean; calculating the squared average of the deviations between the amplitude values and the amplitude mean based on the amplitude mean to obtain the amplitude variance; and selecting the maximum amplitude value from the set of amplitude values as the amplitude peak value.
[0084] In this embodiment, the generation of the multi-domain input field set includes:
[0085] Based on the time-frequency feature set, the spatial coordinate set of measurement points, and the multi-level spatial grid structure, for each coarse and fine grid cell in the multi-level spatial grid structure, the measurement points belonging to the grid cell are determined. Within the grid cell, the time-frequency features of the belonging measurement points in the time-frequency feature set are weighted and normalized according to the preset aggregation rules. The processing results are written into the corresponding grid cell. After all grid cells within the coarse and fine grid ranges are written, a time-frequency feature domain input field aligned with the coarse and fine grids is formed on the multi-level spatial grid structure.
[0086] Based on the spatial grid feature set and multi-level spatial grid structure, for each coarse and fine grid cell in the multi-level spatial grid structure, the vibration energy features and statistical features corresponding to the grid cell are read from the spatial grid feature set. The vibration energy features and statistical features are encoded into the spatial feature vector of the grid cell. The spatial feature vector is written into the corresponding grid cell. After all grid cells in the coarse and fine grid ranges are written, a spatial heat map domain input field aligned with the multi-level spatial grid structure is formed on the multi-level spatial grid structure.
[0087] Based on the scanning parameter set, numerical normalization and encoding processing are performed on the sound source power, sweep frequency excitation signal, playback time, radar focal length, scanning mode, scanning angle, scanning speed, frequency range and amplitude range in the scanning parameter set. The encoded scanning parameters are arranged into scanning parameter vectors in a preset order. In the multi-level spatial grid structure, for each coarse grid and each fine grid cell, the scanning parameter vector is copied into the operator kernel physical modulation vector corresponding to the grid cell. The operator kernel physical modulation vectors of all grid cells are arranged on the multi-level spatial grid structure to generate the scanning parameter modulation field.
[0088] Based on the time-frequency feature domain input field, the spatial heatmap domain input field, and the scanning parameter modulation field, the three input fields are spliced together according to the channel dimension on a multi-level spatial grid structure. The spliced multi-channel grid features are organized into a multi-domain input field set on the multi-level spatial grid structure, and the multi-domain input field set is input into the improved ANO model.
[0089] In this embodiment, the construction and use of the improved ANO model includes:
[0090] Based on a multi-domain input field set and a multi-level spatial grid structure, input channels are set for the time-frequency feature domain input field, spatial heatmap domain input field, and scanning parameter modulation field in the multi-domain input field set on grid nodes aligned with the multi-level spatial grid structure. The three input channels are mapped and aligned on the multi-level spatial grid structure to construct the multi-domain operator input structure of the improved ANO model. A multi-domain attention operator structure is set in the improved ANO model, with the multi-level spatial grid structure as the operator's scope.
[0091] The improved ANO model is an improved form of the Attentive Neural Operators model;
[0092] In the improved ANO model, a multi-domain operator input structure is used as input. A time-frequency operator branch is constructed on the input channel corresponding to the time-frequency feature domain input field, and a spatial operator branch is constructed on the input channel corresponding to the spatial heatmap domain input field. An attention weight calculation unit is set inside the improved ANO model. The attention weights of each grid node in the time-frequency operator branch and the spatial operator branch are jointly modulated based on the operator kernel physical modulation vector generated by the scanning parameter modulation field, so as to obtain the modulated time-frequency operator branch and the modulated spatial operator branch.
[0093] On the coarse and fine grids of the multi-level spatial grid structure, coarse and fine grid operators are configured based on the coarse and fine grid indices, respectively, for the modulated time-frequency operator branch and the modulated spatial operator branch.
[0094] On the coarse grid, coarse grid operator operations are performed on the grid node features of the modulated time-frequency operator branch and the modulated spatial operator branch based on the coarse grid index. At each coarse grid node, the coarse grid operator performs attention-weighted aggregation of the multi-domain features of neighboring coarse grid nodes based on the modulated attention weights, and performs weighted combination of the aggregation results based on the operator kernel at the coarse grid scale to generate coarse grid structural response features representing the structural response state of the coarse grid nodes. The coarse grid structural response features corresponding to all coarse grid nodes are arranged according to the coarse grid index to generate a coarse grid structural response feature field. On the fine grid, the modulated time-frequency operator branch and the spatial operator branch are then processed based on the fine grid index. The grid node features of the modulated spatial operator branch are subjected to fine-grid operator operations. At each fine-grid node, the fine-grid operator performs attention-weighted aggregation on the multi-domain features of the neighboring fine-grid nodes based on the modulated attention weights, and performs feature enhancement operations on the weighted aggregation results based on the operator kernel at the fine-grid scale to generate fine-grid crack local features representing the local response state of the crack. The fine-grid crack local features corresponding to all fine-grid nodes are arranged according to the fine-grid index to generate a fine-grid crack local feature field. The coarse-grid structural response feature field and the fine-grid crack local feature field are used as the feature output results of the improved ANO model for the multi-domain input field set.
[0095] In this embodiment, the generation of the crack risk field includes:
[0096] Based on the coarse and fine grid indices of the multi-level spatial grid structure, a coarse grid index is assigned to each coarse grid node in the response feature field of the coarse grid structure, and a fine grid index is assigned to each fine grid node in the local feature field of the crack in the fine grid. According to the spatial inclusion relationship between the coarse and fine grid indices on the multi-level spatial grid structure, the coarse grid node to which each fine grid node belongs is determined, and a set of coarse and fine grid mapping relationships is generated.
[0097] Based on the set of coarse and fine mesh mapping relationships, the coarse mesh structural response features of each coarse mesh node are read in the coarse mesh structural response feature field, and the fine mesh crack local features of the fine mesh node corresponding to the coarse mesh node are read in the fine mesh crack local feature field. On each fine mesh node, a cross-scale correlation metric is calculated based on the coarse mesh structural response features and the fine mesh crack local features. The cross-scale correlation metric is normalized to generate a cross-scale attention weight set for each fine mesh node.
[0098] A cross-scale attention mechanism is constructed on a multi-level spatial grid structure based on a set of coarse and fine grid mapping relationships and a set of cross-scale attention weights. At each fine grid node, the cross-scale attention weights are used to perform a weighted combination of the coarse grid structure response features of the corresponding coarse grid node and the local features of the fine grid cracks of the fine grid node to generate the crack risk value of the fine grid node. The crack risk values of all fine grid nodes are arranged on the multi-level spatial grid structure according to the fine grid index to generate a crack risk field defined on the multi-level spatial grid structure.
[0099] In this embodiment, the generation of the marking results for the crack areas on the building facade includes:
[0100] Based on the distribution of the crack risk field on the multi-level spatial grid structure, the crack risk value of each grid cell is read from the crack risk field. According to the preset crack risk threshold set, the crack risk value of each grid cell is judged. Grid cells with crack risk values greater than or equal to the first crack risk threshold are marked as crack candidate grid cells, and grid cells with crack risk values less than the first crack risk threshold are marked as non-crack grid cells. The marking results of crack candidate grid cells and non-crack grid cells are organized according to the grid index in the multi-level spatial grid structure to generate a crack candidate marking matrix.
[0101] Based on the crack candidate marker matrix and multi-level spatial grid structure, the spatial adjacency relationship between grid cells is determined according to the coarse grid index and the fine grid index in the multi-level spatial grid structure. Connectivity analysis is performed on the crack candidate grid cells within the coarse grid and the fine grid respectively. The set of crack connected regions composed of spatially adjacent crack candidate grid cells is identified on the multi-level spatial grid structure. For each crack connected region, the number of grid cells in the crack connected region and the cumulative sum of crack risk values of the grid cells in the crack connected region are calculated. Based on the preset threshold for the number of grid cells in the connected region and the preset cumulative crack risk threshold, crack connected regions with a number of grid cells lower than the threshold for the number of grid cells in the connected region and a cumulative crack risk value lower than the cumulative crack risk threshold are removed, thus generating a set of effective crack connected regions.
[0102] Based on the set of effective crack connected regions, grid cells belonging to the same effective crack connected region are aggregated into a single crack region in a multi-level spatial grid structure. In the coordinate system of the building facade, the outer contour extraction operation is performed on the set of fine grid cells in each crack region to generate the corresponding crack region contour.
[0103] For each fine mesh cell within a crack region, the crack area estimation result is calculated based on the number of fine mesh cells and the area parameter of a single fine mesh cell in the building facade coordinate system. A set of crack area estimation results is generated, and the crack region outline and the set of crack area estimation results are stored in correspondence with the set of effective crack connected regions to generate the building facade crack region marking result.
[0104] Based on the crack risk field and the crack area marking results of the building facade, the crack risk values in the crack risk field are color-coded according to the preset color mapping rules in the coordinate system of the building facade, generating a crack risk heat map of the building facade. The crack area outline and crack area estimation results set from the crack area marking results of the building facade are superimposed on the crack risk heat map of the building facade, and a building crack scanning monitoring report containing the crack risk heat map, crack area outline and crack area estimation results is output.
[0105] Example 1:
[0106] To verify the feasibility of this invention in practice, it was applied to a crack scanning and monitoring task on the exterior facade of a concrete building. The building's facade has a large area, with complex factors such as varying degrees of weathering, uneven surface coating thickness, and unstable lighting in some areas. Traditional manual inspections struggle to achieve continuous coverage, and image-based crack identification methods are significantly affected by lighting and surface texture, making it difficult to accurately identify minute or hidden cracks. This invention employs a combined acoustic vibration excitation and laser vibration radar scanning approach. Through multi-domain input field construction and an improved ANO model, combined with a cross-scale attention mechanism, it achieves risk identification and area marking of cracks on the building's exterior facade. This fully demonstrates the invention's ability to solve problems such as the difficulty of large-area scanning, the difficulty of identifying hidden cracks, and the difficulty of separating structural response from crack characterization.
[0107] In practical application, a scanning device consisting of a directional sound source and a laser vibration radar was set up on-site to continuously excite the facade with acoustic vibration and collect the vibration response. The sound source power was set to 150W, the frequency range of the sweep excitation signal was set to 100Hz to 3200Hz, and the amplitude range was set to 30 to 70dB. The radar focal length was set to 12m, the scanning speed was set to 0.5m per second, and the scanning angle covered the entire facade area. During the scanning process, the length of the acoustic vibration response time-series data collected per second was approximately 4096 points, and a total of 31250 acoustic vibration response time-series data points were obtained for the entire facade. Each data point corresponds to a measurement point location, forming a large-scale acoustic vibration response time-series data set.
[0108] Spectral analysis was performed on the acoustic vibration response time-series data set. Frequency domain amplitude characteristics were obtained through Fourier transform, and a time-frequency feature tensor for each measurement point was constructed through time-frequency transformation. Noise estimation and validity assessment were performed on each time-series data, eliminating data with a signal-to-noise ratio below 15dB, resulting in 29,320 valid measurement points. Subsequently, coordinate transformation was performed based on the polar coordinates of the measurement points recorded during the scanning process and the rectangular coordinates of the building image, aligning all measurement points to the unified exterior facade coordinate system of the building. A coarse grid of 20cm and a fine grid of 5cm were created, ultimately generating 2,980 coarse grids and 12,120 fine grids. Each measurement point was assigned to its corresponding grid cell using the spatial coordinate set and the time-frequency feature set. Statistical characteristics such as vibration energy characteristics, mean amplitude, amplitude variance, and peak amplitude of each grid cell were calculated, forming a spatial grid feature set.
[0109] In the construction of the multi-domain input field, the time-frequency feature set is mapped to a time-frequency feature domain input field aligned with the coarse and fine grids, and the spatial grid feature set is mapped to a spatial heatmap domain input field. Simultaneously, the scan parameter set is converted into a physical modulation vector field of the operator kernel, forming a scan parameter modulation field. An improved ANO model is constructed, defining coarse-grid operators and fine-grid operators on the coarse and fine grids respectively. Multi-domain attention operator processing is performed on the multi-domain input field set. The attention weights of the time-frequency operator branch and the spatial operator branch are modulated by the modulation vector generated by the scan parameter modulation field, establishing a physically consistent correlation between the structural response and local crack features.
[0110] During model execution, the coarse-grid operator generates a coarse-grid structural response feature field to describe the overall structural stress response trend; the fine-grid operator generates a fine-grid crack local feature field to characterize the local vibration changes of fine-scale cracks. The system establishes a cross-scale attention mechanism based on a multi-level spatial grid structure, creating a dynamic link between coarse-scale structural information and fine-scale crack characteristics. The modulation effect of the coarse-grid structural response feature field enhances the sensitivity of the fine-grid crack local feature field to anomalous vibration modes, ultimately generating a crack risk field.
[0111] After the crack risk field is generated, the system performs threshold judgment and connectivity analysis on each grid cell to identify and mark crack regions. To verify the effectiveness of the method of the present invention, a comparative experiment was conducted with the traditional image crack recognition method (traditional method A) and the vibration response analysis method based on a single time-frequency feature (traditional method B). The experiments statistically analyzed data such as crack recognition accuracy, fine crack detection rate, false detection rate, and scan integrity index, and the following experimental comparison results were obtained.
[0112] Table 1. Comparison of experimental results between the method of the present invention and the traditional method.
[0113] Indicator Name Traditional Method A Traditional Method B Method of the present invention Crack identification accuracy 82.6% 87.3% 91.4% Detection rate of minute cracks 61.2% 68.5% 81.7% False positive rate 9.8% 8.6% 6.4% Scan integrity 78.1% 84.4% 90.2% Number of hidden cracks detected 17 articles 24 articles 38 items Average processing time per square meter 12.5 seconds 10.8 seconds 11.2 seconds
[0114] As can be seen from the data results in Table 1, the method of the present invention improves the crack identification accuracy by about 8.5 percentage points compared with the traditional method A and by about 4.1 percentage points compared with the traditional method B. This indicates that the present invention can form a more stable crack identification basis and improve the overall identification accuracy in the cross-scale fusion process of the response feature field of coarse grid structure and the local feature field of fine grid crack.
[0115] Regarding the detection rate of fine cracks, the method of this invention improves by 20.5 percentage points compared with the traditional method A and by 13.2 percentage points compared with the traditional method B, indicating that the cross-scale attention mechanism enhances the characteristics of fine-scale abnormal vibrations, enabling the model to have a stronger ability to distinguish between fine cracks and early cracks.
[0116] Regarding the false detection rate, the false detection rate of the method of this invention is 6.4%, which is lower than that of traditional method A (9.8%) and traditional method B (8.6%), representing a reduction of 3.4 percentage points and 2.2 percentage points respectively. This indicates that the physical modulation information introduced by the multi-domain input field construction enables the model to more accurately distinguish noise vibrations generated in non-crack regions, thereby effectively reducing false detections.
[0117] In terms of scan completeness, the method of this invention achieves 90.2%, which is higher than the 78.1% of traditional method A and the 84.4% of traditional method B, representing improvements of 12.1 percentage points and 5.8 percentage points, respectively. This indicates that the multi-level spatial grid structure can effectively reduce area omissions and improve the overall scan coverage in large-area facade scanning tasks.
[0118] Regarding the ability to detect hidden cracks, the method of this invention detected 38 hidden cracks, an increase of 21 compared to the 17 detected by the traditional method A, and an increase of 14 compared to the 24 detected by the traditional method B. This indicates that the multi-level spatial feature coupling mechanism enables the model to capture the internal response differences of the structure even when faced with complex situations such as surface texture being occluded, uneven lighting, or surface material coverage, thereby improving the ability to detect hidden cracks.
[0119] In terms of average processing time, the method of this invention is about 11.2 seconds per square meter, which is slightly higher than the traditional method B, but better than the traditional method A. This shows that the processing efficiency remains stable even with the addition of multi-domain input field calculation and cross-scale attention fusion, indicating that the present invention still maintains a good balance between computational load and recognition performance.
[0120] Example 2:
[0121] To verify the feasibility of this invention in practice, it was applied to the task of scanning and monitoring exterior cracks in a building with mortar and paint as the exterior wall material and a building with ceramic tiles as the exterior wall material, and on-site test examples were conducted.
[0122] The equipment used for this external wall hollowness inspection was the YSV-100 scanning laser vibration meter. This device is a laser Doppler vibration measurement device with functions such as automatic scanning, distance measurement, arbitrary grid setting, and automatic supplementation of invalid data. It has a built-in laser vibration measurement module, optical camera, and optical pan-tilt unit, and features small size, long measurement distance, high accuracy, high efficiency, and strong anti-interference ability.
[0123] This inspection employed a combination of acoustic excitation and laser detection to scan for hollow areas in the exterior wall. A directional sound source emitted a broadband audio frequency (500-4000 Hz, customizable) towards the target wall surface, controlling the sound intensity at approximately 80 dB. This caused forced vibrations in the wall, the frequency spectrum and energy of which depended on the wall's material and condition. If the wall was uniformly solid, the amplitude of the forced vibrations would be small, with energy concentrated in the low-frequency range below 1000 Hz. Conversely, if there were hollow areas within the wall, the forced vibrations would be significantly amplified, with a marked increase in high-frequency energy. Furthermore, the larger the area and depth of the hollow areas, the stronger the forced vibration energy. Therefore, comparing the energy changes of the forced vibrations could effectively identify the presence of hollow areas within the exterior wall. Due to the limited energy of acoustic waves and the high stiffness of the wall, the amplitude of the forced vibrations was very small, typically within tens of nanometers. To accurately detect this signal, laser vibrometric technology was employed, with a measurement sensitivity reaching the 10 picometer level.
[0124] In the task of scanning and monitoring cracks on the exterior facade of buildings with mortar and paint as the exterior wall material:
[0125] First test: Test parameters: Distance from the device to the wall: 4.97m; Measurement point interval: 15×15cm; Test area: 1.42m² 2 Number of measurement points: 63; Scanning time: 65s; Hollow rate: 30.2%; Detection results are as follows: Figure 4 , Figure 5 The dark-colored measuring points in the figure are significant hollow areas, and most of the measuring points have been manually verified to be correct.
[0126] Second test: Test parameters: Distance of equipment from wall: 4.97m; Measurement point interval: 15×15cm; Test area: 1.42m² 2 Number of measurement points: 63; Scanning time: 90s; Hollow rate: 36.5%; Detection results are as follows: Figure 6 , Figure 7The dark-colored measuring points in the figure are significant hollow areas, and most of the measuring points have been manually verified to be correct.
[0127] In the task of scanning and monitoring cracks on the exterior facade of buildings with ceramic tile as the exterior wall material:
[0128] First test: Test parameters: Distance of equipment from wall: 9.28m; Measurement point interval: 20×20cm; Test area: 7.6m² 2 Number of measurement points: 190; Scanning time: 176s; Hollow rate: 43.7%; Detection results are as follows: Figure 8 , Figure 9 , Figure 10 The dark-colored measuring points in the figure are significant hollow areas, and most of the measuring points have been manually verified to be correct.
[0129] Second test: Test parameters: Distance from the device to the wall: 16.26m; Measurement point interval: 20×20cm; Test area: 5.12m²; Number of measurement points: 128; Scanning time: 103s; Hollow rate: 0; Test results are as follows: Figure 11 , Figure 12 The wall surface showed no abnormalities; the two measuring points marked with boxes were false detection points. Figure 12 The second and third measuring points on the lower left were misdetected because workers nearby were using electric picks to break up the concrete ground, which introduced a large noise interference when the equipment was collecting vibration signals, causing the equipment to misjudge that there was a large amount of vibration energy at this location.
[0130] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A deep learning-based method for scanning and monitoring building cracks, characterized in that, Includes the following steps: A swept-frequency acoustic vibration excitation signal is radiated onto the exterior of the building, and the acoustic vibration response time-series data of the exterior of the building are collected to construct an acoustic vibration response time-series data set; Spectral analysis is performed on the acoustic vibration response time-series data set, and a time-frequency feature tensor is constructed based on the spectral analysis results to generate a time-frequency feature set; Based on the measurement point location records of the acoustic vibration response time series data set, a multi-level spatial grid structure is constructed, each measurement point is assigned to a grid cell in the multi-level spatial grid structure, the vibration energy characteristics and statistical characteristics of each grid cell are calculated, and a spatial grid feature set is generated. The time-frequency feature set is mapped to the time-frequency feature domain input field, and the spatial grid feature set is mapped to the spatial heat map domain input field, which are then combined into a multi-domain input field set. An improved ANO model is constructed, using a multi-level spatial grid structure as the operator's domain, performing multi-domain attention operator processing, defining coarse grid operators and fine grid operators, and generating a coarse grid structure response feature field and a fine grid crack local feature field. A cross-scale attention mechanism is established based on a multi-level spatial grid structure. The response feature field of the coarse grid structure is used to modulate the local feature field of the crack in the fine grid to generate a crack risk field. Based on the crack risk field, threshold judgment and connectivity analysis are performed on the crack risk value corresponding to each grid cell in the crack risk field to generate the marking results of crack areas on the building facade.
2. The method for scanning and monitoring building cracks based on deep learning according to claim 1, characterized in that, The construction of the acoustic vibration response time-series data set includes: Set up a subset of sound source parameters, a subset of radar parameters, and a subset of scanning control parameters. Perform numerical range setting, step resolution setting, and unit unification processing on the three subsets respectively to generate a set of scanning parameters. Based on the geometric contour information and scanning mode of the building facade, the scanning angle and scanning speed are obtained from the subset of scanning control parameters. A preset scanning path covering the detection area is generated in the coordinate system of the building facade. The scanning trajectory point sequence and corresponding time sequence of the laser vibration radar are determined to form the preset scanning path data. Based on the set of scanning parameters and preset scanning path data, the directional sound source is controlled to radiate frequency-sweeping acoustic and vibration excitation signals to the exterior of the building, and the laser vibration radar is controlled to collect acoustic and vibration response time-series data corresponding to the time series in sequence, generating the original acoustic and vibration response time-series data. The original acoustic and vibration response time series data were synchronized and the measurement points were numbered to construct an acoustic and vibration response time series data set.
3. The method for scanning and monitoring building cracks based on deep learning according to claim 1, characterized in that, The generation of the time-frequency feature set includes: Perform spectral analysis on the acoustic and vibration response time-series data set, including Fourier transform, time-frequency transform, noise estimation, validity assessment, and power spectrum calculation; The spectral analysis results are reorganized and stacked according to the measurement point dimension, time dimension, and frequency dimension to construct a time-frequency feature tensor; The data segments corresponding to each measurement point in the time-frequency feature tensor are extracted into measurement point-level feature units and organized to form a time-frequency feature set.
4. The method for scanning and monitoring building cracks based on deep learning according to claim 1, characterized in that, The generation of the spatial grid feature set includes: Based on the measurement point location records in the acoustic vibration response time series data set, the polar coordinate measurement point coordinates output by the laser vibration radar are paired one by one with the rectangular coordinates of the building facade image. The polar coordinate measurement point coordinates are then transformed to obtain the spatial coordinates of the measurement points, and a set of spatial coordinates of the measurement points is generated. Under the building facade coordinate system corresponding to the spatial coordinate set of measurement points, a coarse grid and a fine grid are divided according to a preset grid size. Coarse grid index and fine grid index are set respectively to construct a multi-level spatial grid structure covering the building facade. The time-frequency features corresponding to each measuring point in the time-frequency feature set are assigned to the corresponding coarse and fine grid cells according to the position of the measuring point's spatial coordinates in the multi-level spatial grid structure. Based on the time-frequency features corresponding to the measuring points assigned to the grid cells, the vibration energy features and statistical features of the grid cells are calculated, and a spatial grid feature set is generated.
5. The method for scanning and monitoring building cracks based on deep learning according to claim 1, characterized in that, The generation of the multi-domain input field set includes: Based on the time-frequency feature set, the spatial coordinate set of measurement points, and the multi-level spatial grid structure, for each coarse and fine grid cell in the multi-level spatial grid structure, the measurement points belonging to the grid cell are determined. The time-frequency features of the belonging measurement points in the time-frequency feature set are weighted and normalized according to the preset aggregation rules. The processing results are written into the corresponding grid cell. After all grid cells are written, the time-frequency feature domain input field is formed. In a multi-level spatial grid structure, the vibration energy characteristics and statistical characteristics of the grid cell are read from the spatial grid feature set, encoded into the spatial feature vector of the grid cell, and written into the corresponding grid cell. After all grid cells are written, the spatial heat map domain input field is formed. Numerical normalization and encoding are performed on the set of scanning parameters. The encoded scanning parameters are arranged into a scanning parameter vector. In the grid cells of the multi-level spatial grid structure, the scanning parameter vector is copied into the operator kernel physical modulation vector corresponding to the grid cell and arranged to generate a scanning parameter modulation field. Based on the time-frequency feature domain input field, the spatial heat map domain input field, and the scanning parameter modulation field, the three input fields are spliced together and organized into a multi-domain input field set.
6. The method for scanning and monitoring building cracks based on deep learning according to claim 1, characterized in that, The construction and use of the improved ANO model include: Based on a multi-domain input field set and a multi-level spatial grid structure, input channels are set for the multi-domain input field set on grid nodes aligned with the multi-level spatial grid structure. The three input channels are mapped and aligned on the multi-level spatial grid structure to construct the multi-domain operator input structure of the improved ANO model. A multi-domain attention operator structure is set, with the multi-level spatial grid structure as the operator scope. In the improved ANO model, a multi-domain operator input structure is used as input. A time-frequency operator branch is constructed on the input channel corresponding to the time-frequency feature domain input field, and a spatial operator branch is constructed on the input channel corresponding to the spatial heatmap domain input field. Attention weight calculation units are set up, and the attention weights of each grid node in the time-frequency operator branch and the spatial operator branch are jointly modulated based on the operator kernel physical modulation vector. On the coarse and fine grids of the multi-level spatial grid structure, coarse and fine grid operators are configured based on the coarse and fine grid indices, respectively, for the modulated time-frequency operator branch and spatial operator branch; Perform coarse mesh operator operations to generate a coarse mesh structural response feature field; perform fine mesh operator operations to generate a fine mesh crack local feature field.
7. The method for scanning and monitoring building cracks based on deep learning according to claim 1, characterized in that, The generation of the crack risk field includes: In the coarse mesh structure response feature field, a coarse mesh index is assigned to each coarse mesh node, and in the fine mesh crack local feature field, a fine mesh index is assigned to each fine mesh node. Based on the spatial inclusion relationship between the coarse mesh index and the fine mesh index, the coarse mesh node to which each fine mesh node belongs is determined, and a set of coarse and fine mesh mapping relationships is generated. Based on the set of coarse and fine mesh mapping relationships, the coarse mesh structural response characteristics of each coarse mesh node and the local characteristics of fine mesh cracks of the corresponding fine mesh node are read. Cross-scale correlation measures are calculated, normalization is performed, and a cross-scale attention weight set is generated. A cross-scale attention mechanism is constructed on a multi-level spatial grid structure. At each fine grid node, the cross-scale attention weight is used to perform a weighted combination of the coarse grid structure response features of the corresponding coarse grid node and the local features of the fine grid cracks of the fine grid node, thereby generating crack risk values for the fine grid nodes and arranging them to generate a crack risk field.
8. The method for scanning and monitoring building cracks based on deep learning according to claim 1, characterized in that, The generation of the marking results for the crack areas on the building facade includes: Based on the distribution of the crack risk field on a multi-level spatial grid structure, the crack risk value of each grid cell is read from the crack risk field, a threshold judgment is performed, crack candidate grid cells and non-crack grid cells are marked, the marking results of crack candidate grid cells and non-crack grid cells are organized, and a crack candidate marking matrix is generated. Based on the crack candidate marker matrix and multi-level spatial grid structure, the spatial adjacency relationship between grid cells is determined, connectivity analysis is performed on the crack candidate grid cells, the set of crack connected regions composed of spatially adjacent crack candidate grid cells is identified, the number of grid cells and the cumulative sum of crack risk values in the crack connected regions are calculated, and sets with the number of grid cells and the cumulative sum of crack risk values below the threshold are removed to generate a set of effective crack connected regions. Based on the set of effective crack connected regions, the mesh cells belonging to the same effective crack connected region are aggregated into a single crack region. The outer contour extraction operation is performed on the set of fine mesh cells in each crack region to generate the corresponding crack region contour. For each crack region, calculate the crack area estimation result, generate a crack area estimation result set, and store the crack region outline and crack area estimation result set corresponding to the effective crack connected region set to generate the building facade crack region marking result; The crack risk values in the crack risk field are color-coded to generate a crack risk heat map of the building facade. The crack area outline and crack area estimation results are overlaid and displayed, and a building crack scanning monitoring report containing the crack risk heat map, crack area outline and crack area estimation results is output.