Building crack intelligent detection and prediction method based on multi-modal data fusion
By collecting multi-source heterogeneous data through drone multi-sensors and combining generative adversarial networks and multimodal fusion networks, the problems of all-round monitoring and accurate prediction of building crack detection are solved, and intelligent safety monitoring and early warning of building structures are realized.
Patent Information
- Application Number
- CN202510648222.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-12
AI Technical Summary
Existing building crack detection methods mainly rely on manual visual inspection, which makes it difficult to achieve all-round and real-time monitoring. In addition, the data collected by a single sensor cannot fully describe the state of the building structure, resulting in inaccurate detection results and large deviations in prediction results, and unable to provide reliable early warnings.
Unmanned aerial vehicles (UAVs) equipped with multiple sensors are used to acquire multi-source heterogeneous data. Image training data is enhanced through generative adversarial networks. A dual-branch multimodal fusion network is designed to extract features. LSTM and structural mechanics models are used to predict crack trends, achieving intelligent monitoring and early warning.
It realizes all-round and all-weather building structure monitoring, improves the integrity and accuracy of data, enhances the adaptability and robustness of the model to complex environments, improves the reliability of crack detection and the accuracy of prediction, and provides reliable early warning for building structure safety.
Smart Images

Figure CN120635638A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of crack detection technology, and in particular to a method for intelligent detection and prediction of building cracks based on multimodal data fusion. Background Art
[0002] In the field of construction engineering, the safety of building structures is always of paramount importance. As a key factor affecting structural safety, the accurate detection and development trend prediction of building cracks are of immeasurable value in ensuring the stability and durability of buildings.
[0003] Existing methods for detecting building cracks primarily rely on manual visual inspection. Inspectors use simple tools like magnifying glasses and crack calipers to closely observe cracks on the building surface. This method is limited by the inspector's field of view and inspection angle. Inspectors are unable to directly observe difficult-to-reach areas of the building structure, such as the tops of high-rise building facades and the internal structures of large bridges, making it easy to miss important crack information. Furthermore, manual inspections struggle to achieve comprehensive, real-time monitoring of building structures and fail to promptly detect rapidly developing cracks. Furthermore, the accuracy of manual inspection results is significantly influenced by the subjective factors of the inspector. Different inspectors have different criteria for judging crack width, length, morphology, and other characteristics, making it difficult to ensure the reliability of the inspection data.
[0004] In terms of data acquisition technology, existing technologies often use a single type of sensor for data collection. For example, some detection methods rely solely on optical cameras to capture image data from the building surface. While these methods can visually demonstrate the appearance of cracks, they fail to capture deeper information such as temperature changes and stress distribution within the structure. Using only vibration sensors, on the other hand, can only monitor the structural vibration response and insufficiently reflect the direct characteristics of cracks. The limitations of this single sensor data collection prevent the information obtained from comprehensively describing the state of the building structure, making it difficult to accurately analyze the root causes and development trends of cracks. In practical applications, buildings operate in complex and diverse environments, such as varying light intensity, weather conditions, and ambient interference, resulting in varying data quality. Existing data processing methods struggle to effectively enhance data from these complex environments, resulting in poor adaptability of detection models trained on these data. These models lack robustness when used with data from diverse environments, making them prone to misjudgments and missed detections. For crack development trend prediction, simple statistical models or empirical formulas have been used. These methods fail to fully consider the mechanical properties of the building structure and the nonlinear factors involved in crack development. When a building structure is subjected to actual stress, crack expansion is not only related to the current load but also to a variety of factors, including the structure's historical stress conditions and material properties. Simple prediction models cannot accurately capture these complex relationships, resulting in significant discrepancies between predicted results and actual crack development. This inability to provide reliable early warnings for building structural safety can delay the implementation of appropriate reinforcement and repair measures, posing potential risks to building structural safety.
[0005] In summary, the existing building crack detection and prediction technology has many problems that need to be solved in terms of detection methods, data collection and processing, and prediction accuracy. An innovative and comprehensive technical solution is urgently needed to overcome these difficulties and improve the level of building structure safety monitoring. Summary of the Invention
[0006] The present invention aims to solve one of the problems existing in the background technology.
[0007] To this end, the present invention provides a method for intelligent detection and prediction of building cracks based on multimodal data fusion. It constructs a multi-source heterogeneous dataset by using multiple sensors equipped on drones, enhances image training data through GAN, extracts features using a dual-branch multimodal fusion network, and predicts crack trends based on LSTM and structural mechanics models. Finally, the integrated system is deployed to achieve intelligent monitoring and early warning of building structural safety.
[0008] The technical solution adopted by the present invention to solve its technical problem is:
[0009] A method for intelligent detection and prediction of building cracks based on multimodal data fusion, comprising:
[0010] Step 1: Obtain high-resolution images, infrared thermal imaging data, and structural vibration sensor data of the building structure, and construct a multi-source heterogeneous dataset, including image datasets, thermal imaging datasets, and vibration datasets;
[0011] Step 2: Generate adversarial networks (GANs) to simulate complex environmental interference and dynamically enhance image training data based on actual environmental interference conditions.
[0012] Step 3: Design a dual-branch multimodal fusion network. Design a visual feature extraction branch to extract edge and texture features of cracks using a convolutional neural network. Design a multimodal fusion branch to align and weightedly fuse visual, thermal, and vibration data using a cross-modal attention layer.
[0013] Step 4: Based on the long short-term memory network (LSTM) and the structural mechanics model, develop a crack trend prediction module to provide early warning for building structural safety.
[0014] Furthermore, in step 2, the generative adversarial network is composed of a generator G and a discriminator D, the generator G is represented by a function G(z), which inputs a random noise vector z and outputs simulated interference data of the random noise vector z; the input data x, the probability D(x) of the discriminator D that the output data x is the real data, D(x)∈[0,1].
[0015] Furthermore, in step 2, the image training data is dynamically enhanced and the enhanced image is x' i =F(x i , G1(z1), G2(z2), G3(z3), G4(z4)), where x i is the original image of the i-th training sample, N is the total number of training samples, G1, G2, G3, G4 are multiple generators, which are used to simulate the interference factors of illumination change, rain fog, dust, and dynamic shadow, and output the interference feature map G respectively. j (z j ), where z j is the input noise vector of the j-th interference generator, and F is the fusion network.
[0016] Furthermore, the fusion network F respectively processes the original image x i And each interference data G j (z j ) to extract convolutional features and generate high-dimensional feature representation; dynamic weights α are assigned to different interference features i , so that the fusion process can highlight the environmental interference information that is most helpful for model training; the extracted high-dimensional features are globally averaged pooled to obtain the scalar value s of the i-th high-dimensional feature j, j is the high-dimensional feature index, generating the dynamic weight of the i-th high-dimensional feature
[0017] Furthermore, in step 3, the process of extracting the edge and texture features of the crack by the visual feature extraction branch is as follows: the convolution kernel of the visual feature extraction branch slides on the image, performs a convolution operation with the local area of the image, extracts the local features of the image, generates a feature map, and then performs a maximum pooling operation. After processing by multiple convolution layers and maximum pooling layers, the fully connected layer further processes the flattened feature vector F, integrates the previously extracted features, and finally obtains a feature vector that can accurately characterize the crack morphology.
[0018] Furthermore, the visual feature extraction branch calculation process is: in, is the output feature value of the lth convolutional layer at position (i, j); is the value of the convolution kernel of the lth layer at position (m, n, c), M and N correspond to the height and width of the convolution kernel respectively, C l-1 Indicates the number of channels in the previous layer, b l is the bias of the lth layer, I is the enhanced image data, i+m and j+n are used to determine the position of the image in the height and width directions, and c represents the image channel.
[0019] Furthermore, the maximum pooling operation: in, is the output feature value of the lth pooling layer at position (i, j), s×s is the pooling kernel, is the output feature value of the lth convolutional layer at position (s×i+m,s×j+n), and m and n are used to traverse each element in the pooling window.
[0020] Furthermore, the output O of the fully connected layer is calculated by the following formula: O = σ(W f ·F+b f ), where the weight matrix of the fully connected layer is W f , bias is b f , σ is the ReLU activation function.
[0021] Furthermore, in the multimodal fusion branch, the vibration feature vector Z, the visual feature vector V processed by the visual feature extraction branch, and the thermal imaging feature vector T are normalized to Z′, V′, and T′ respectively, and then the cross-modal attention score is calculated to obtain the visual feature and thermal imaging feature attention score A vt , vibration feature and thermal imaging feature attention score A zt , vibration feature and visual feature attention score A vz, according to the weighted coefficient of each attention score, the final fused multimodal feature vector F fasion .
[0022] Furthermore, in step 4, a basic framework is built based on the long short-term memory network LSTM and the structural mechanics model, and the input historical crack data sequence is X = {x1, x2, ..., x T}, where x T Represents the multimodal feature vector F extracted at time step t fasion ; The input gate of the long short-term memory network is represented by i t , the calculation formula is i t =σ(W ii x t +W hi h t-1 +b i ), where W ii and W hi are the input weight matrix and the hidden layer weight matrix, b i is the bias term, σ is the sigmoid activation function, h t-1 is the hidden state of the previous time step; the LSTM output gate outputs the crack length of the future time step; the LSTM network is used to predict the crack development trend and the potential risk is assessed in combination with the structural mechanics model: the crack length predicted by the LSTM network in the future time step is substituted back into the structural mechanics model to calculate the crack extension driving force G; if the predicted crack length increases, the calculated G is closer to or exceeds the critical fracture toughness G of the material. c , it indicates that there is a high potential risk for the future structure, providing an accurate early warning for the safety of building structures.
[0023] The beneficial effects of the present invention are:
[0024] 1. This invention utilizes drones equipped with multiple sensors to comprehensively collect high-resolution images, infrared thermal imaging data, and structural vibration sensor data from building structures, constructing a multi-source, heterogeneous dataset. Compared to traditional single-source data collection methods, the information obtained is more comprehensive, reflecting the building structure's condition from multiple dimensions. This significantly improves data integrity and accuracy, providing a rich and reliable data foundation for subsequent crack detection and prediction.
[0025] 2. This invention uses a generative adversarial network (GAN) to dynamically enhance image training data, simulating complex environmental interference. This enables the model to accurately identify and analyze crack characteristics even in complex environments such as varying lighting and occlusion. Compared to traditional data processing methods, this significantly improves the model's adaptability and robustness to complex environments, reduces false positives and missed detections due to environmental factors, and improves the reliability and stability of crack detection.
[0026] 3. This invention employs a dual-branch multimodal fusion network. The visual feature extraction branch utilizes a convolutional neural network to effectively extract visual features such as crack edges and textures. The multimodal fusion branch uses a cross-modal attention layer to perform feature alignment and weighted fusion of visual, thermal, and vibration data. This multimodal fusion approach fully integrates the advantages of different data types. Compared to single-modal data processing, it can more comprehensively and deeply mine crack-related information, thereby improving the ability to characterize crack characteristics and laying a solid foundation for accurate crack detection and trend prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The present invention will be further described below with reference to the accompanying drawings and examples.
[0028] Figure 1 It is a flowchart of the intelligent detection and prediction method of building cracks based on multimodal data fusion in the present invention.
[0029] Figure 2 It is an implementation flow chart of the intelligent detection and prediction method of building cracks based on multimodal data fusion in the present invention. DETAILED DESCRIPTION
[0030] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0031] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, features defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.
[0032] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0033] An intelligent detection and prediction method for building cracks based on multimodal data fusion,
[0034] Step 1: Construct a multi-source heterogeneous dataset. Use a drone equipped with a high-resolution camera, infrared thermal imager, and structural vibration sensor to collect all-round data on the building structure. Obtain high-resolution images, infrared thermal imaging data, and structural vibration sensor data, and construct a multi-source heterogeneous dataset, including image datasets, thermal imaging datasets, and vibration datasets. Specifically, dataset collection includes the following steps:
[0035] Step 1.1 Data acquisition equipment preparation:
[0036] The drone is equipped with a high-resolution camera to capture detailed information on the surface of the building structure and obtain high-quality image data; it is also equipped with an infrared thermal imager to detect the temperature distribution on the surface of the building structure and obtain infrared thermal imaging data; it is equipped with a structural vibration sensor, which uses the homography transformation method to eliminate the influence of wind and its own vibration to obtain structural vibration sensor data.
[0037] Step 1.2 Comprehensive data collection:
[0038] Plan the drone's flight path to cover all parts of the building structure, including the roof, walls, beams and columns, which are key structural components. The flight path is designed based on the shape, height and complexity of the building. For buildings with regular shapes, rectangular or spiral flight paths are used.
[0039] Step 1.3 Associate the calibration dataset:
[0040] During the data collection process, the drone's positioning system is used to record the drone's location information, and the collected data is spatially positioned in combination with the geographic coordinates and relative position relationship of the building structure; ensuring that the data collected by different devices can accurately correspond in space. For a building structure at a specific location, its high-resolution image, infrared thermal imaging data and structural vibration data are accurately associated according to the spatial coordinates and timestamps to ensure the correspondence of the data; the collected data is calibrated, and through timestamp matching and spatial coordinate conversion, the data inconsistency caused by the difference in equipment collection time and location is eliminated to ensure the consistency of multi-source data.
[0041] Step 1.4: Build a multi-source heterogeneous dataset:
[0042] The collected high-resolution images are screened and preprocessed to remove blurry, noisy, and incomplete images; images that meet the requirements are organized into image datasets and classified and labeled according to different building structure parts, crack types, etc. The labeled information should include the location, size, and shape characteristics of the cracks; in the case of infrared thermal imaging data, the temperature distribution data is converted into visual thermal imaging images, and the processed thermal imaging images are organized into thermal imaging datasets and classified and labeled as well. The labeled information includes the location and temperature value of the temperature anomaly area; when processing structural vibration sensor data, the characteristic parameters of vibration, such as vibration frequency, amplitude, and phase, are extracted; the processed vibration data is organized into a vibration dataset and classified and labeled. The labeled information includes the location and characteristic parameter values of the vibration; the image dataset, thermal imaging dataset, and vibration dataset are integrated together to construct a multi-source heterogeneous dataset;
[0043] Step 2: Generate adversarial networks (GANs) to simulate complex environmental interference, improve the model's adaptability and robustness to complex environments, dynamically enhance image training data based on actual environmental interference conditions, and adjust data enhancement strategies to ensure the effectiveness of model training.
[0044] Step 2.1: Simulate complex environment interference based on generative adversarial network:
[0045] The generative adversarial network consists of a generator G and a discriminator D. Specifically:
[0046] When simulating complex environmental interference, the goal of the generator G is to generate simulated interference data similar to the real interference data; let the input random noise vector be z, and represent the generator G as a function G(z), and the function output is the simulated interference data generated with the random noise vector z as input.
[0047] The role of the discriminator D is to determine whether the input data is real interference data or simulated interference data generated by the generator; the output of the discriminator D for the input data x is a probability value D(x)∈[0,1], which represents the probability that the input data is real data; when training GAN, the parameters of the generator and discriminator are optimized through the minimax game; the goal of the generator is to maximize the probability that the discriminator misclassifies its generated data as real data, while the goal of the discriminator is to maximize the probability of correctly classifying real data and generated data. The process is described by the following optimization objective function:
[0048]
[0049] Among them, V(D,G) is the optimization objective function of GAN, x~P data (x) is a random variable x subject to P data The probability distribution of (x), P data (x) is the distribution of real interference data, p z (z) is the distribution of the noise vector z, and E represents the expected operation; through continuous iterative training, the generator gradually generates realistic complex environment interference data, and the generator is trained as a generator that simulates lighting interference and a generator that simulates occlusion interference.
[0050] Step 2.2 Dynamically enhance the image training data:
[0051] Let the original training data be X = {x1, x2, ..., x N}, where x i Represents the i-th training sample, and N is the total number of training samples. To enhance the robustness of the model to various complex environmental interferences, a multi-interference generator fusion strategy is proposed. Multiple generators G1, G2, G3, and G4 are trained separately to simulate the interference factors of illumination changes, rain, fog, dust, and dynamic shadows, and output interference feature maps G respectively. j (z j ), where z j is the input noise vector of the j-th interference generator.
[0052] In order to achieve the composite enhancement effect of interference, the fusion network F is used to perform nonlinear fusion on each interference feature map to generate the final enhanced image x' i , the formula is as follows:
[0053] x' i =F(x i , G1(z1), G2(z2), G3(z3), G4(z4))
[0054] Among them, the F module processes the input original image and each interference feature map as follows:
[0055] Step 2.2.1: For the original image x i And each interference data G j (z j ) performs convolutional feature extraction to generate high-dimensional feature representation.
[0056] Step 2.2.2 Assign dynamic weights α to different interference features i , so that the environmental interference information that is most helpful for model training can be highlighted during the fusion process.
[0057] Perform global average pooling on the high-dimensional features extracted in step 2.2.1 to obtain the scalar value s j , j is the high-dimensional feature index, generating dynamic weight α i :
[0058]
[0059] Among them, α i represents the dynamic weight of the i-th high-dimensional feature, s i is the scalar value of the i-th high-dimensional feature.
[0060] Step 2.2.3: The weighted high-dimensional features are subjected to nonlinear activation and convolution operations to reconstruct an enhanced image x' with multiple real environment interference effects. i .
[0061] Step 2.3: Adjust the data enhancement strategy based on the actual environmental interference situation:
[0062] To ensure the effectiveness of model training, the data enhancement strategy is adjusted in real time according to the actual environmental interference conditions. By monitoring the changes in interference factors in the real environment, environmental sensors are used to collect light intensity and temperature data, and the occlusion conditions around the actual building structures are manually assessed regularly. The interference vector {e1, e2, ..., e m}, where m is the number of interference factors, e m Represents the measured value of the mth interference factor; the environmental interference vector is mapped to the input noise vector z to adjust the output of the generator G, so that the simulated interference data generated by the generator is more consistent with the interference conditions in the actual environment, ensuring that the model can fully adapt to the changes in the actual complex environment during training, and improving the adaptability and robustness of the model.
[0063] Step 3: Design a dual-branch multimodal fusion network, design a visual feature extraction branch, and use a convolutional neural network to extract the edge and texture features of the crack; design a multimodal fusion branch, and use a cross-modal attention layer to align and weight the visual, thermal imaging, and vibration data. The fusion feature output flow chart is as follows: Figure 2 As shown;
[0064] Step 3.1 Design visual feature extraction branch:
[0065] Let the image data after the enhancement in step 2 be I, and the calculation formula of the visual feature extraction branch is as follows:
[0066]
[0067] in, is the output feature value of the lth convolutional layer at position (i, j); is the value of the convolution kernel of the lth layer at position (m, n, c), M and N correspond to the height and width of the convolution kernel respectively, C l-1 Indicates the number of channels in the previous layer, b l is the bias of the lth layer, I is the enhanced image data, i+m and j+n are used to determine the position of the image in the height and width directions, and c represents the image channel. Through this formula, the convolution kernel slides on the image, performs a convolution operation with the local area of the image, extracts the local features of the image, generates a feature map, and then performs a maximum pooling operation:
[0068]
[0069] in, is the output feature value of the lth pooling layer at position (i, j), s×s is the pooling kernel, is the output eigenvalue of the lth convolutional layer at position (s×i+m,s×j+n), m and n are used to traverse each element in the pooling window; after processing by multiple convolutional layers and pooling layers, the fully connected layer further processes the flattened eigenvector F; the weight matrix of the fully connected layer is W f , bias is b f , then the output O of the fully connected layer is calculated by the following formula:
[0070] O=σ(W f ·F+b f )
[0071] Among them, σ is the ReLU activation function. After the operation of the fully connected layer, the previously extracted features are integrated to finally obtain a feature vector that can accurately characterize the crack morphology.
[0072] Step 3.2 Design multimodal fusion branch:
[0073] Let the vibration eigenvector be Z, the visual eigenvector and thermal imaging eigenvector output from step 3.1 be V and T respectively. First, normalize the modal data. The eigenvectors after preprocessing are V′, T′, and Z′ respectively:
[0074] V′=norm(V)
[0075] T′=norm(T)
[0076] Z′=norm(Z)
[0077] Among them, norm(·) is the normalization function, and then the cross-modal attention score, the visual feature and thermal imaging feature attention score A are calculated. vt The dot product operation yields:
[0078]
[0079] Among them, d v is the dimension of the feature vector. By calculating the dot product of two feature vectors and normalizing them, we can get the similarity score between them. Similarly, we can calculate A vz and A t2 , respectively, represent the attention scores between vision and vibration, thermal imaging and vibration; the weighting coefficient is calculated according to the attention score, where the weighting coefficient α of the thermal imaging feature is t The calculation formula can be expressed as:
[0080]
[0081] Among them, exp(·) is the exponential function, α t Indicates the relative importance weight of thermal imaging features in the fusion process; similarly, the visual feature weighting coefficient α can be calculated v and vibration characteristic weighting coefficient α z ; The final fused multimodal feature vector F fasion Calculated by the following formula:
[0082] F fasion =α v ·V′+α t ·T′+α z ·Z′
[0083] Through the above-mentioned weighted summation method, the feature vectors of different modes are fused according to their respective importance weights to obtain a comprehensive multi-modal feature vector.
[0084] Step 4: Develop a crack trend prediction module based on the long short-term memory network (LSTM) and the structural mechanics model. This module uses the LSTM network to predict crack development trends and combines it with the structural mechanics model to assess potential risks and provide early warning for building structural safety.
[0085] Step 4.1: Build a basic framework based on the long short-term memory network (LSTM) and the structural mechanics model:
[0086] Assume that the input historical crack data sequence is X={x1,x2,...,x T}, where x TRepresents the multimodal feature vector f extracted in step 3 at time step t fasion , T is the length of the time series; the input gate of the long short-term memory network is represented by i t , and its calculation formula is:
[0087] i t =σ(W ii x t +W hi h t-1 +b i )
[0088] Among them, W ii and W hi are the input weight matrix and the hidden layer weight matrix, b i is the bias term, σ is the sigmoid activation function, h t-1 is the hidden state of the previous time step; LSTM processes the long-term dependencies in the historical crack data, and the output gate outputs the crack length in the future time step; the crack extension driving force G is related to the stress intensity factor K and is calculated by the following formula:
[0089]
[0090] Where E is the elastic modulus of the material, K is the stress intensity factor. For a rectangular beam with a single-side crack, when subjected to a uniformly distributed load q, the calculation formula for the stress intensity factor K is:
[0091]
[0092] Where a is the crack length and h is the height of the beam. The potential risk of crack expansion is analyzed from a mechanical perspective based on the actual parameters and stress conditions of the building structure.
[0093] Step 4.2: Use the LSTM network to predict the crack development trend and combine it with the structural mechanics model to assess potential risks: The LSTM network predicts the crack length in the future time step by learning the pattern in the historical feature vector; and then substitutes the predicted crack length into the structural mechanics model to calculate the crack extension driving force G. If the predicted crack length increases, the calculated G will be closer to or exceed the critical fracture toughness G of the material. c , it indicates that there is a high potential risk for the future structure, providing an accurate early warning for the safety of building structures.
[0094] Step 5: System integration and deployment: Integrate the proposed method into the intelligent detection and prediction system for building cracks, deploy the system to actual application scenarios, and provide intelligent monitoring and early warning services for building structural safety.
[0095] With the above-described preferred embodiments of the present invention as inspiration, and with reference to the above description, relevant personnel are fully capable of making various changes and modifications without departing from the technical spirit of this invention. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. A method for intelligent detection and prediction of building cracks based on multimodal data fusion, characterized in that: include: Step 1: Obtain high-resolution images, infrared thermal imaging data, and structural vibration sensor data of the building structure, and construct a multi-source heterogeneous dataset, including image datasets, thermal imaging datasets, and vibration datasets; Step 2: Generate adversarial networks (GANs) to simulate complex environmental interference and dynamically enhance image training data based on actual environmental interference conditions. Step 3: Design a dual-branch multimodal fusion network, design a visual feature extraction branch, and use a convolutional neural network to extract the edge and texture features of the crack; Design a multimodal fusion branch to align and weightedly fuse visual, thermal imaging, and vibration data through a cross-modal attention layer. Step 4: Based on the long short-term memory network (LSTM) and the structural mechanics model, develop a crack trend prediction module to provide early warning for building structural safety.
2. The method for intelligent detection and prediction of building cracks based on multimodal data fusion according to claim 1 is characterized in that: In step 2, the generative adversarial network consists of a generator G and a discriminator D. The generator G is represented by a function G(z). The random noise vector z is input, and the function outputs simulated interference data of the random noise vector z. The input data x, the probability D(x) of the output data x of the discriminator D is the real data, D(x)∈[0,1].
3. The intelligent detection and prediction method for building cracks based on multimodal data fusion according to claim 2 is characterized in that: In step 2, the image training data is dynamically enhanced and the enhanced image is x' i =F(x i , G1(z1), G2(z2), G3(z3), G4(z4)), where x i is the original image of the i-th training sample, N is the total number of training samples, G1, G2, G3, G4 are multiple generators, which are used to simulate the interference factors of illumination change, rain fog, dust, and dynamic shadow, and output the interference feature map G respectively. j (z j ), where z j is the input noise vector of the j-th interference generator, and F is the fusion network.
4. The method for intelligent detection and prediction of building cracks based on multimodal data fusion according to claim 3 is characterized in that: The fusion network F respectively processes the original image x i And each interference data G j (z j ) to extract convolutional features and generate high-dimensional feature representation; dynamic weights α are assigned to different interference features i , so that the fusion process can highlight the environmental interference information that is most helpful for model training; the extracted high-dimensional features are globally averaged pooled to obtain the scalar value s of the i-th high-dimensional feature j , j is the high-dimensional feature index, generating the dynamic weight of the i-th high-dimensional feature 5. The intelligent detection and prediction method for building cracks based on multimodal data fusion according to claim 1 is characterized in that: In step 3, the process of extracting the edge and texture features of the crack by the visual feature extraction branch is as follows: the convolution kernel of the visual feature extraction branch slides on the image, performs a convolution operation with the local area of the image, extracts the local features of the image, generates a feature map, and then performs a maximum pooling operation. After processing by multiple convolution layers and maximum pooling layers, the fully connected layer further processes the flattened feature vector F, integrates the previously extracted features, and finally obtains a feature vector that can accurately characterize the crack morphology.
6. The intelligent detection and prediction method for building cracks based on multimodal data fusion according to claim 5 is characterized in that: The calculation process of the visual feature extraction branch is: in, is the output feature value of the lth convolutional layer at position (i, j); is the value of the convolution kernel of the lth layer at position (m, n, c), M and N correspond to the height and width of the convolution kernel respectively, C l-1 Indicates the number of channels in the previous layer, b l is the bias of the lth layer, I is the enhanced image data, i+m and j+n are used to determine the position of the image in the height and width directions, and c represents the image channel.
7. The intelligent detection and prediction method for building cracks based on multimodal data fusion according to claim 5 is characterized in that: Max pooling operation: in, is the output feature value of the lth pooling layer at position (i, j), s×s is the pooling kernel, is the output feature value of the lth convolutional layer at position (s×i+m,s×j+n), and m and n are used to traverse each element in the pooling window.
8. The intelligent detection and prediction method for building cracks based on multimodal data fusion according to claim 5 is characterized in that: The output O of the fully connected layer is calculated by the following formula: O = σ(W f ·F+b f ), where the weight matrix of the fully connected layer is W f , bias is b f , σ is the ReLU activation function.
9. The intelligent detection and prediction method for building cracks based on multimodal data fusion according to claim 5 is characterized in that: In the multimodal fusion branch, the vibration feature vector Z, the visual feature vector V and the thermal imaging feature vector T processed by the visual feature extraction branch are normalized to Z′, V′ and T′ respectively, and then the cross-modal attention score is calculated to obtain the visual feature and thermal imaging feature attention score A. vt , vibration feature and thermal imaging feature attention score A zt , vibration feature and visual feature attention score A vz , according to the weighted coefficient of each attention score, the final fused multimodal feature vector F fasion .
10. The intelligent detection and prediction method for building cracks based on multimodal data fusion according to claim 1 is characterized in that: In step 4, a basic framework is built based on the long short-term memory network LSTM and the structural mechanics model, and the input historical crack data sequence is X = {x1, x2, ..., x T }, where x T Represents the multimodal feature vector F extracted at time step t fasion ; The input gate of the long short-term memory network is represented by i t , the calculation formula is i t =σ(W ii x t +W hi h t-1 +b i ), where W ii and W hi are the input weight matrix and the hidden layer weight matrix, b i is the bias term, σ is the sigmoid activation function, h t-1 is the hidden state of the previous time step; the LSTM output gate outputs the crack length of the future time step; the LSTM network is used to predict the crack development trend and the potential risk is assessed in combination with the structural mechanics model: the crack length predicted by the LSTM network in the future time step is substituted back into the structural mechanics model to calculate the crack extension driving force G; if the predicted crack length increases, the calculated G is closer to or exceeds the critical fracture toughness G of the material. c , it indicates that there is a high potential risk for the future structure, providing an accurate early warning for the safety of building structures.
Citation Information
Cited By
Thermal vibration heterogeneous data fusion conveyor belt edge tearing detection method
CN120986946A
A thermal-vibration heterogeneous data fusion conveyor belt edge tear detection method
CN120986946B
Urban environment performance prediction method and system based on multi-modal fusion and attention enhancement
CN121053470A
Method and system for monitoring running state of ultra-high-efficiency dust explosion-proof brake motor
CN121276323A
Concrete crack development situation prediction method and system based on time sequence image
CN121304681A