Concrete crack detection method and system based on multi-modal sensing and deep learning

By combining multimodal sensing and deep learning methods with SH array ultrasound, distributed optical fiber and millimeter-wave radar, the limitations of existing technologies for concrete crack detection have been solved, achieving high-precision, full-dimensional crack detection and digital characterization.

CN122066640APending Publication Date: 2026-05-19CHINA RAILWAY BIYUAN WATER SERVICE KUNMING CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA RAILWAY BIYUAN WATER SERVICE KUNMING CO LTD
Filing Date
2025-12-30
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies for detecting internal cracks in concrete have limitations due to their reliance on single detection methods. They are difficult to achieve full-angle coverage, have high construction costs, limited ability to detect deep cracks, and are difficult to unify and comprehensively interpret different detection results, making it impossible to construct a complete three-dimensional crack model.

Method used

A multimodal sensing and deep learning approach is adopted, utilizing a multimodal sensing module composed of SH array ultrasonic units, distributed optical fiber units, and millimeter-wave radar units, combined with a GPS synchronization clock and data fusion correction module, to establish a unified timestamp and three-dimensional spatial coordinate system. Feature extraction and fusion are performed through an improved Transformer network to generate crack features.

Benefits of technology

It achieves high-precision, multi-dimensional crack detection, generating digital crack profiles that include crack size parameters, spatial coordinates, and predicted probabilities. This improves the accuracy and anti-interference capability of the detection, overcomes the shortcomings of single sensors, and ensures the stability of the system in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122066640A_ABST
    Figure CN122066640A_ABST
Patent Text Reader

Abstract

The invention provides a concrete crack detection method and system based on multi-modal sensing and deep learning, and the method comprises the steps: inputting collected multi-source heterogeneous data to a data fusion correction module, and carrying out the processing through time synchronization, space alignment, feature fusion and environment compensation, and obtaining the corrected crack feature information; the crack three-dimensional modeling module constructs and outputs a three-dimensional model of the concrete internal crack and characteristic parameters of the three-dimensional model based on the output corrected crack characteristics; a unified time reference and a space coordinate system are established through a space-time synchronization mechanism of a GPS synchronous clock and Ethernet communication, alignment of SH ultrasonic, optical fiber strain and millimeter wave radar data under the same space-time reference system is achieved, and joint expression of multi-modal features in a unified semantic space is achieved on the basis. And meanwhile, the influence of environmental factors on a detection result is effectively reduced in combination with a dynamic environment compensation mechanism, so that high-precision and stable detection of the internal cracks of the concrete is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of concrete crack detection technology, and more specifically, to a concrete crack detection method and system based on multimodal sensing and deep learning. Background Technology

[0002] The current methods for detecting internal cracks in concrete mainly rely on the following technologies: (1) Ultrasonic detection: Cracks inside concrete are located by sound wave reflection. Shear horizontal waves (SH waves) are used as detection signals. SH waves have high sensitivity to the tilt angle (0°–90°) of cracks when propagating in concrete, and are especially suitable for detecting vertical and inclined cracks. Ultrasonic signals are transmitted and received by an array of SH wave probes arranged in a ring, and the original signals of all transmit-receive pairs are recorded by full matrix capture (FMC) technology. (2) Fiber optic sensor detection: The strain distribution of concrete is monitored by single-mode fiber, and strain anomalies can be detected. (3) Millimeter-wave radar detection: High-frequency electromagnetic waves are emitted by millimeter-wave radar, and surface and shallow cracks are detected by receiving reflected signals. Synthetic aperture radar (SAR) technology is used to enhance signal processing and improve imaging resolution.

[0003] The existing single detection methods have limitations: (1) Ultrasonic detection is limited by the probe arrangement and beam directionality, making it difficult to cover cracks at all angles (0°–90°). Traditional ultrasonic imaging has a high rate of missed detection for thin cracks. (2) Distributed fiber optic sensors are pre-deployed or pre-embedded, resulting in high construction costs. Moreover, their detection results are mainly based on strain time-series information, which cannot directly form spatial imaging results of cracks. It is difficult to construct a three-dimensional crack model and can only provide local strain gradient information. (3) Millimeter-wave radar detection has limited ability to detect deep cracks (such as cracks in the walls of nuclear power plants), and the reflected signal attenuation is significant. Therefore, the common objective problem of the existing technologies is that each detection technology usually operates in a single mode and independently. There is a lack of a unified time reference, a unified spatial coordinate system, and a cross-modal semantic association mechanism, which makes it difficult to effectively align and comprehensively interpret different detection results. Consequently, it is difficult to obtain complete spatial distribution information of cracks inside concrete and to construct a reliable three-dimensional crack model. Summary of the Invention

[0004] In view of this, in order to solve the problems mentioned in the background technology, a concrete crack detection method and system based on multimodal sensing and deep learning is proposed.

[0005] The objective of this invention can be achieved through the following technical solutions:

[0006] This invention provides a method and system for detecting concrete cracks based on multimodal sensing and deep learning, comprising:

[0007] S10: Perform multimodal data acquisition and processing operations, and use the multimodal sensing module (10) composed of SH array ultrasonic unit (11), distributed optical fiber unit (12) and millimeter-wave radar unit (13) to detect the concrete structure; wherein, the SH array ultrasonic unit (11) acquires crack image data inside the concrete and records it as ultrasonic image data, the distributed optical fiber unit (12) acquires strain time series data on the concrete surface, and the millimeter-wave radar unit (13) acquires crack image data on the concrete surface and shallow layer and records it as radar image data;

[0008] S20: Input the multi-source heterogeneous data collected in step S10 into the data fusion and correction module (20) for processing; this step specifically includes:

[0009] S21: Provide a unified timestamp for the multimodal sensing module (10) by synchronizing the GPS clock, and establish a unified three-dimensional spatial coordinate system based on the pre-calibrated physical coordinates of each sensor, so as to map the crack image data, strain time series data and crack image data into the same spatiotemporal reference system;

[0010] S22: Perform environmental factor correction on the data before fusion. Based on the temperature, humidity and vibration data collected by environmental sensors, perform environmental factor compensation correction on the multimodal data to reduce the impact of environmental changes on the crack characteristics.

[0011] S23: An improved Transformer network is used to extract and fuse features from spatiotemporally aligned multimodal data to generate fused crack features.

[0012] Preferably, as one possible implementation, step S10 specifically includes:

[0013] S101: Including the pre-scanning stage, firstly, the millimeter-wave radar unit (13) is used to quickly scan the concrete surface and mark the suspicious crack areas on the surface and in the shallow layer;

[0014] S102: Includes the execution of a precise detection phase, controlling the SH array ultrasonic unit (11) to perform full-focus imaging on the marked suspicious area, while the distributed optical fiber unit (12) monitors the strain in real time; if the optical fiber strain data triggers a preset abnormal threshold, the high-precision detection mode of the SH array ultrasonic unit (11) is activated.

[0015] Preferably, as one possible implementation, the spatiotemporal alignment in step S21 specifically includes:

[0016] S211: Time synchronization, the GPS synchronization clock module provides a unified UTC timestamp for the data acquisition actions of the SH array ultrasonic unit (11), distributed optical fiber unit (12) and millimeter wave radar unit (13), so that the data of each sensor carries the same time reference;

[0017] S212: Spatial synchronization. Before detection, the physical coordinates of each sensor are recorded and a unified three-dimensional coordinate system (X,Y,Z) is established. The scanning data coordinates of the millimeter-wave radar, the layout coordinates of the fiber optic strain monitoring points, and the geometric layout coordinates of the SH ultrasonic array are transmitted to the edge computing terminal via Ethernet protocol and projected onto the same spatial grid to achieve precise spatial overlap of different modal data.

[0018] Preferably, as an implementation scheme; in step S23, the improved Transformer network utilizes a cross-modal attention mechanism to enable acoustic image features from the SH array ultrasonic unit (11), electromagnetic image features from the millimeter-wave radar unit (13), and strain timing features from the distributed optical fiber unit (12) to interact, constrain, and suppress noise in the same semantic space, thereby extracting consistent crack features across physical mechanisms.

[0019] A concrete internal crack detection system based on multimodal sensing and deep learning, comprising:

[0020] The multimodal sensing module (10) includes an SH array ultrasonic unit (11), a distributed optical fiber unit (12), and a millimeter-wave radar unit (13), which is used to acquire multi-source heterogeneous detection data of concrete structures. The SH array ultrasonic unit (11) is used to acquire crack image data inside the concrete and record it as ultrasonic image data. The distributed optical fiber unit (12) is used to acquire strain time series data of the concrete surface. The millimeter-wave radar unit (13) is used to acquire crack image data of the concrete surface and shallow layers and record it as radar image data.

[0021] A data fusion correction module (20) is communicatively connected to the multimodal sensing module (10) and is used to process the multi-source heterogeneous detection data; the data fusion correction module (20) includes:

[0022] The spatiotemporal alignment submodule (21) is used to provide a unified timestamp for the multimodal sensing module (10) through a GPS synchronization clock, and to establish a unified three-dimensional spatial coordinate system based on the pre-calibrated physical coordinates of each sensor, so as to map the ultrasonic image data, strain time series data and radar image data into the same spatiotemporal reference system.

[0023] The environmental compensation and correction submodule (22) is used to use the temperature, humidity and vibration data collected by the environmental sensor to correct the drift between the ultrasonic image data and the strain time series data through the Gaussian process regression model, and to eliminate the interference of vibration on the radar image data through the band-stop filter to obtain the corrected data.

[0024] The deep learning fusion submodule (23) is used to extract and fuse features of the spatiotemporally aligned multimodal data using an improved Transformer network. The improved Transformer network includes a spatial attention module to process ultrasonic image data and radar image data, and a temporal attention module to process fiber strain time series features, thereby generating fused crack features.

[0025] The output module (30) is connected to the data fusion correction module (20) and is used to generate and output a three-dimensional model of the crack and its feature parameters based on the corrected crack features.

[0026] Preferably, as one possible implementation; the SH array ultrasound unit (11) includes:

[0027] The shear wave transmitter is arranged in a ring, with an operating frequency of 100-500 kHz and an array element spacing of 10 cm, covering a detection grid of 0.5 m × 0.5 m; the signal processing unit is configured to acquire signals using full matrix acquisition technology and apply a full focusing method to convert the raw data into tomographic images.

[0028] Preferably, as one possible implementation, the distributed optical fiber unit (12) uses flexible packaging technology to directly attach single-mode optical fiber to the concrete surface, and its data processing unit is configured to trigger abnormal area marking by setting a strain gradient threshold in order to determine the location of cracks.

[0029] Preferably, as one possible implementation, the millimeter-wave radar unit (13) is a millimeter-wave synthetic aperture radar operating in the 60 GHz band, and its data processing unit is configured to perform a back-projection imaging algorithm on the acquired radar echo signal to generate a distribution image of the concrete surface and shallow cracks.

[0030] Preferably, as one possible implementation, the spatiotemporal alignment submodule (21) further includes a networked integration unit for connecting the demodulator of the distributed optical fiber unit (12), the millimeter-wave radar unit (13) and the edge computing terminal via Ethernet to form a high-speed data communication network for real-time transmission of sensor data with spatiotemporal coordinates.

[0031] Preferably, as one possible implementation, the environmental compensation correction submodule (23) includes:

[0032] The temperature and humidity compensation unit is configured to use a Gaussian process regression model to dynamically correct the ultrasonic velocity and fiber strain values ​​based on real-time collected environmental temperature and humidity data.

[0033] The vibration filtering unit is configured to use an accelerometer to monitor vibration and use a band-stop filter to eliminate interference frequency bands to radar signals.

[0034] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects:

[0035] Analysis of the above technical solution shows that the method includes: S10: performing multimodal data processing operation, using a multimodal sensing module (10) composed of SH array ultrasonic unit (11), distributed optical fiber unit (12) and millimeter-wave radar unit (13) to detect concrete structure; wherein, the SH array ultrasonic unit (11) acquires ultrasonic image data inside the concrete and records it as ultrasonic image data, the distributed optical fiber unit (12) acquires strain time series data of concrete surface, and the millimeter-wave radar unit (13) acquires crack image data of concrete surface and shallow layer and records it as radar image data;

[0036] S20: Input the multi-source heterogeneous data collected in step S10 into the data fusion correction module (20) for processing; this step specifically includes: S21: Provide a unified timestamp for the multimodal sensing module (10) through GPS synchronization clock, and establish a unified three-dimensional spatial coordinate system based on the pre-calibrated physical coordinates of each sensor, and map the ultrasonic image data, strain time series data and radar image data into the same spatiotemporal reference system; S22: Perform environmental factor correction on the data before fusion, use the temperature, humidity and vibration data collected by environmental sensors, correct the drift of ultrasonic image data and strain time series data through Gaussian process regression model, and eliminate the interference of vibration on radar image data through band-stop filter, and then re-fuse to obtain the corrected crack features; S23: Use an improved Transformer network to extract and fuse the spatiotemporally aligned multimodal data, the improved Transformer network includes a spatial attention module to process ultrasonic image data and radar image data, and a temporal attention module to process fiber strain time series features, thereby generating the fused crack features;

[0037] S30: Three-dimensional crack model generation step, based on the corrected crack features output in step S20, constructs and outputs a three-dimensional model of the internal cracks in the concrete and its feature parameters; the feature parameters include crack size parameters.

[0038] This invention provides a concrete crack detection system based on multimodal sensing and deep learning. It integrates complementary data from three physical mechanisms: ultrasound (internal cracks), fiber optics (surface strain time series), and millimeter-wave radar (surface and shallow cracks). This overcomes the limitations of single sensors, such as susceptibility to interference and limited detection dimensions, significantly improving the accuracy and anti-interference capability of crack identification (especially internal and shallow cracks). Simultaneously, it utilizes full-dimensional quantization to characterize multi-source data of cracks. It not only detects the presence of cracks but also generates a "digital crack profile" containing dimensional parameters (penetration, width), three-dimensional spatial coordinates, and predicted probabilities, achieving a transition from qualitative detection to quantitative and model-based diagnosis. Through environmental factor correction modules (temperature and humidity compensation, vibration compensation) and cross-modal attention mechanisms (noise suppression, feature consistency constraints), the system effectively suppresses the impact of complex on-site environments on the data, ensuring system stability in real-world engineering scenarios and thus exhibiting stronger environmental robustness.

[0039] This invention provides a concrete crack detection method and system based on multimodal sensing and deep learning. Its core is to achieve high-precision, full-dimensional, and high-reliability digital detection and characterization of concrete cracks through multi-source heterogeneous data fusion and deep learning intelligent analysis. Attached Figure Description

[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 A schematic diagram of the main process of the concrete crack detection method based on multimodal sensing and deep learning provided by the present invention;

[0042] Figure 2 This is a schematic diagram illustrating the specific operation process of the concrete crack detection method based on multimodal sensing and deep learning provided by the present invention.

[0043] Figure 3 A schematic diagram illustrating the spatiotemporal alignment process of the concrete crack detection method based on multimodal sensing and deep learning provided by this invention.

[0044] Figure 4 The schematic diagram of the concrete crack detection system based on multimodal sensing and deep learning provided by this invention. Detailed Implementation

[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] Example 1

[0047] See Figure 1 This invention provides a concrete crack detection method based on multimodal sensing and deep learning, comprising:

[0048] S10: Perform multimodal data acquisition and processing operations, and use the multimodal sensing module (10) composed of SH array ultrasonic unit (11), distributed optical fiber unit (12) and millimeter-wave radar unit (13) to detect the concrete structure; wherein, the SH array ultrasonic unit (11) acquires crack image data inside the concrete and records it as ultrasonic image data, the distributed optical fiber unit (12) acquires strain time series data on the concrete surface, and the millimeter-wave radar unit (13) acquires crack image data on the concrete surface and shallow layer and records it as radar image data;

[0049] S20: Input the multi-source heterogeneous data collected in step S10 into the data fusion and correction module (20) for processing; this step specifically includes:

[0050] S21: Provide a unified timestamp for the multimodal sensing module (10) by synchronizing the GPS clock, and establish a unified three-dimensional spatial coordinate system based on the pre-calibrated physical coordinates of each sensor, so as to map the ultrasonic image data, strain time series data and radar image data into the same spatiotemporal reference system;

[0051] S22: Perform environmental factor correction on the data before fusion. Use temperature, humidity and vibration data collected by environmental sensors to correct the drift between ultrasonic image data and strain time series data through Gaussian process regression model. Then, use a band-stop filter to eliminate the interference of vibration on radar image data. Finally, re-fuse the data to obtain the corrected crack features.

[0052] S23: An improved Transformer network is used to extract and fuse features from the spatiotemporally aligned multimodal data. The improved Transformer network includes a spatial attention module to process ultrasonic image data and radar image data, and a temporal attention module to process fiber strain time series features, thereby generating fused crack features.

[0053] S30: Three-dimensional crack model generation step. Based on the corrected crack features output in step S20, a three-dimensional model of the internal cracks in the concrete and its feature parameters are constructed and output. The feature parameters include crack size parameters, three-dimensional spatial coordinates (i.e. crack location coordinates) and crack prediction probability.

[0054] It should be noted that in step S20, the data is derived from unified and corrected multimodal data (ultrasonic images, radar images, and fiber optic strain time series) through a deep learning fusion module (especially the cross-modal attention mechanism) and an environmental compensation and correction module. Finally, in step S30, a three-dimensional spatial model of the crack is constructed, and the multidimensional feature parameters of the crack are extracted and output simultaneously to form a complete "digital crack archive".

[0055] Among them, the above-mentioned crack size parameters may include the following aspects, for example: (1) Crack penetration: used to characterize the vertical depth information of the crack extending from the surface to the interior. This is one of the core advantages of the present invention. By fusing and analyzing the internal crack image data obtained by the SH array ultrasonic unit and the surface strain time series data obtained by the distributed optical fiber unit, it is possible to more accurately determine whether the crack is a surface crack, a shallow crack, a deep crack, or even a through crack. (2) Crack width: used to characterize the average width or maximum opening of the crack. Millimeter-wave radar has high sensitivity to the opening of surface cracks, while the abnormal strain gradient collected by the distributed optical fiber unit can indirectly reflect the degree of crack opening. After the two are fused, a reliable estimate of the crack width can be achieved.

[0056] The three-dimensional spatial coordinates (i.e., crack location coordinates) are used to characterize the precise position of the crack's center point or start and end points in a unified three-dimensional coordinate system (X, Y, Z). For example, (X: 2.35m, Y: 5.67m, Z: -0.18m) indicates that the crack is located 18 centimeters below the initial origin.

[0057] The crack prediction probability is a quantitative confidence index generated by a deep learning fusion submodule (usually the output layer of a classification or segmentation network). It characterizes the degree of confidence that a crack exists or belongs to a specific type. For example, a higher value (closer to 1) indicates that the model is more certain that a crack exists at that location (or that the crack belongs to a specific type / severity level), and the basis for this judgment (data features from ultrasound, radar, and fiber optics) is very clear and consistent in the feature space. A lower value (closer to 0) indicates that the model believes there is no crack at that location, or that the features are very vague or contradictory. If the calculated output value is in the middle (e.g., 0.5-0.7), it indicates that there is uncertainty in the model's judgment.

[0058] During step S30, a 3D crack model is generated to output a visualized and quantifiable "digital crack archive," providing a comprehensive basis for structural health assessment. Input: The corrected fusion features output from S20, which already contain comprehensive evidence of cracks in multimodal data. 3D Model Construction: Utilizing the spatial information (from aligned ultrasonic and radar data) contained in the fusion features, the morphology, orientation, and distribution of cracks in 3D space are reconstructed. Feature Parameter Extraction: Crack penetration is accurately determined by fusing ultrasonic (internal depth) and optical fiber (surface strain correlation). For example, if ultrasonic anomalies are detected at depth, and the surface optical fiber shows tension strain at the vertical projection position, it can be inferred to be a deep or penetrating crack. Crack width is estimated by combining radar (directly measuring surface openings) and optical fiber (strain gradient indirectly reflecting width changes), which is more reliable than single-sensor estimation. 3D spatial coordinates are directly derived from the unified 3D coordinate system established in S21 and used throughout; positioning accuracy is determined by the original accuracy of each sensor and the spatial alignment accuracy. Finally, the crack prediction probability is output: This is a comprehensive confidence index output by the fusion network (such as the Softmax layer). It reflects the consistency of multimodal evidence: consistent and strong evidence indicates a high probability; contradictory or weak evidence indicates a low probability. This provides engineers with important decision-making references (e.g., high-probability cracks need to be dealt with immediately, medium-probability cracks need to be reviewed, and low-probability cracks can be observed).

[0059] In summary, this invention provides a concrete crack detection method based on multimodal sensing and deep learning. It integrates complementary data from three physical mechanisms: ultrasound (internal cracks), fiber optics (surface strain time series), and millimeter-wave radar (surface and shallow cracks). This overcomes the limitations of single sensors, such as susceptibility to interference and limited detection dimensions, significantly improving the accuracy and anti-interference capability of crack identification (especially internal and shallow cracks). Furthermore, it utilizes full-dimensional quantization to characterize multi-source data of cracks. It not only detects the presence of cracks but also generates a "digital crack profile" containing dimensional parameters (penetration, width), three-dimensional spatial coordinates, and predicted probabilities, achieving a transition from qualitative detection to quantitative and model-based diagnosis. Through environmental factor correction modules (temperature and humidity compensation, vibration compensation) and cross-modal attention mechanisms (noise suppression, feature consistency constraints), the influence of complex on-site environments on the data is effectively suppressed, ensuring the system's stability in real-world engineering scenarios and thus exhibiting stronger environmental robustness.

[0060] This invention provides a concrete crack detection system based on multimodal sensing and deep learning. Its core is to achieve high-precision, full-dimensional, and high-reliability digital detection and characterization of concrete cracks through multi-source heterogeneous data fusion and deep learning intelligent analysis.

[0061] See Figure 2 The S10 step specifically includes:

[0062] S101: Including the pre-scanning stage, firstly, the millimeter-wave radar unit (13) is used to quickly scan the concrete surface and mark the suspicious crack areas on the surface and in the shallow layer;

[0063] S102: Includes the execution of a precise detection phase, controlling the SH array ultrasonic unit (11) to perform full-focus imaging on the marked suspicious area, while the distributed optical fiber unit (12) monitors the strain in real time; if the optical fiber strain data triggers a preset abnormal threshold, the high-precision detection mode of the SH array ultrasonic unit (11) is activated.

[0064] In the above technical solution, S101-S102 perform a two-stage detection operation, achieving efficient general survey of large-scale structures and high-precision detection of key areas, thus improving the efficiency of large-scale detection. First, radar is used to quickly scan the entire field to locate "suspicious areas," avoiding the time-consuming point-by-point scanning of ultrasound. Then, ultrasonic and fiber optic resources are concentrated to perform high-precision detection on these key areas. Real-time fiber optic monitoring acts as a "trigger," immediately activating the high-precision ultrasonic mode upon detecting abnormal strain, achieving adaptive detection based on structural response, and capable of capturing sudden or expanding cracks.

[0065] See Figure 3 In step S21, spatiotemporal alignment specifically includes:

[0066] S211: Time synchronization, the GPS synchronization clock module provides a unified UTC timestamp for the data acquisition actions of the SH array ultrasonic unit (11), distributed optical fiber unit (12) and millimeter wave radar unit (13), so that the data of each sensor carries the same time reference;

[0067] S212: Spatial synchronization. Before detection, the physical coordinates of each sensor are recorded and a unified three-dimensional coordinate system (X,Y,Z) is established. The scanning data coordinates of the millimeter-wave radar, the layout coordinates of the fiber optic strain monitoring points, and the geometric layout coordinates of the SH ultrasonic array are transmitted to the edge computing terminal via Ethernet protocol and projected onto the same spatial grid to achieve precise spatial overlap of different modal data.

[0068] In the above technical solution, S211-S212 performs detailed spatiotemporal alignment, which is used to explicitly use UTC timestamps and Ethernet to transmit coordinates. It provides a specific and operable implementation plan, ensuring global consistency of time synchronization and real-time and accuracy of spatial data transmission. It is the fundamental guarantee for all subsequent fusion and positioning work.

[0069] Preferably, as an implementation scheme; in step S22, the improved Transformer network utilizes a cross-modal attention mechanism to enable acoustic image features from the SH array ultrasonic unit (11), electromagnetic image features from the millimeter-wave radar unit (13), and strain timing features from the distributed optical fiber unit (12) to interact, constrain, and suppress noise in the same semantic space, thereby extracting consistent crack features across physical mechanisms.

[0070] Example 2

[0071] See Figure 4 Embodiment 2 of the present invention provides a concrete internal crack detection system based on multimodal sensing and deep learning, comprising:

[0072] The multimodal sensing module (10) includes an SH array ultrasonic unit (11), a distributed optical fiber unit (12), and a millimeter-wave radar unit (13), used to acquire multi-source heterogeneous detection data of concrete structures; the SH array ultrasonic unit (11) is used to acquire crack image data inside the concrete and record it as ultrasonic image data; the distributed optical fiber unit (12) is used to acquire strain time series data of the concrete surface; and the millimeter-wave radar unit (13) is used to acquire crack image data of the concrete surface and shallow layers and record it as radar image data.

[0073] A data fusion correction module (20) is communicatively connected to the multimodal sensing module (10) and is used to process the multi-source heterogeneous detection data; the data fusion correction module (20) includes:

[0074] The spatiotemporal alignment submodule (21) is used to provide a unified timestamp for the multimodal sensing module (10) through GPS synchronization clock, and to establish a unified three-dimensional spatial coordinate system based on the pre-calibrated physical coordinates of each sensor, so as to map the crack image data, strain time series data and crack image data into the same spatiotemporal reference system.

[0075] The environmental compensation and correction submodule (22) is used to perform environmental factor correction on the data before fusion. It uses temperature, humidity and vibration data collected by environmental sensors to correct the drift between ultrasonic image data and strain time series data through Gaussian process regression model, and eliminates the interference of vibration on radar image data through band-stop filter. Then it is re-fused to obtain the corrected crack features.

[0076] The deep learning fusion submodule (23) is used to extract and fuse features of the spatiotemporally aligned multimodal data using an improved Transformer network. The improved Transformer network includes a spatial attention module to process ultrasonic image data and radar image data, and a temporal attention module to process fiber strain time series features, thereby generating fused crack features.

[0077] The output module (30) is connected to the data fusion correction module (20) and is used to generate and output a three-dimensional model of the crack and its feature parameters based on the corrected crack features.

[0078] Preferably, as one possible implementation; the SH array ultrasound unit (11) includes:

[0079] The shear wave transmitter is arranged in a ring, with an operating frequency of 100-500 kHz and an array element spacing of 10 cm, covering a detection grid of 0.5 m × 0.5 m; the signal processing unit is configured to acquire signals using full matrix acquisition technology and apply a full focusing method to convert the raw data into tomographic images.

[0080] Regarding the acquisition of internal tomographic images by the SH array ultrasound unit, the following operations are included: the SH array ultrasound unit uses shear horizontal wave transmitters arranged in a ring to cover the detection area; the SH array ultrasound unit eliminates detection blind spots and acquires all transmit and receive signals under the ring arrangement.

[0081] Preferably, as one possible implementation, the distributed optical fiber unit (12) uses flexible packaging technology to directly attach single-mode optical fiber to the concrete surface, and its data processing unit is configured to trigger abnormal area marking by setting a strain gradient threshold in order to determine the location of cracks.

[0082] Preferably, as one possible implementation, the millimeter-wave radar unit (13) is a synthetic aperture radar operating in the 60 GHz band, configured to generate a surface crack distribution image using a back projection algorithm.

[0083] Preferably, as one possible implementation, the spatiotemporal alignment submodule (21) further includes a networked integration unit for connecting the demodulator of the distributed optical fiber unit (12), the millimeter-wave radar unit (13) and the edge computing terminal via Ethernet to form a high-speed data communication network for real-time transmission of sensor data with spatiotemporal coordinates.

[0084] Preferably, as one possible implementation, the environmental compensation correction submodule (22) includes:

[0085] The temperature and humidity compensation unit is configured to use a Gaussian process regression model to dynamically correct the ultrasonic velocity and fiber strain values ​​based on real-time collected environmental temperature and humidity data.

[0086] The vibration filtering unit is configured to use an accelerometer to monitor vibration and use a band-stop filter to eliminate interference frequency bands to radar signals.

[0087] In the above technical solutions, sound velocity correction based on temperature and humidity environmental factors is an important technical means. Specifically, increased temperature leads to a decrease in the elastic modulus of concrete, resulting in a decrease in sound velocity; increased humidity increases the internal moisture content of concrete, increasing its density and slightly increasing the sound velocity. By quantifying the influence of temperature and humidity on sound velocity, the positioning accuracy of ultrasonic testing can be significantly improved, making it suitable for detecting concrete cracks in complex environments such as bridges and dams.

[0088] It should be noted that the concrete crack detection system based on multimodal sensing and deep learning provided in this embodiment of the invention includes: the system consists of a multimodal sensing module, a data fusion and correction module, and an output module. The multimodal sensing module includes an SH array ultrasonic unit, a distributed optical fiber unit, and a millimeter-wave radar unit. The SH array ultrasonic unit acquires images of internal cracks, the distributed optical fiber unit acquires strain time-series data, and the millimeter-wave radar unit acquires images of surface and shallow cracks.

[0089] The multimodal sensing module performs a pre-scanning phase by scanning the surface with a millimeter-wave radar unit to mark suspicious areas, and a precise detection phase by generating crack images of the marked areas with an SH array ultrasonic unit. Simultaneously, a distributed fiber optic unit monitors strain to trigger abnormal area marking. The data fusion and correction module sequentially executes a spatiotemporal alignment submodule, a deep learning fusion submodule, and an environmental compensation and correction submodule. The spatiotemporal alignment submodule achieves temporal and spatial synchronization of multimodal data. The deep learning fusion submodule uses an improved Transformer architecture to perform cross-modal feature fusion on the spatiotemporally aligned data to generate a three-dimensional crack model. The environmental compensation and correction submodule performs environmental factor correction on the data before fusion. The output module outputs the three-dimensional crack model and its feature parameters.

[0090] Regarding the monitoring of strain-triggered abnormal area marking, the process includes: the millimeter-wave radar unit uses synthetic aperture radar mode to scan the concrete surface to generate a surface crack distribution map and mark suspicious areas; the SH array ultrasonic unit uses full-matrix capture technology to acquire transmit and receive signals for the marked suspicious areas during the precision detection phase; the SH array ultrasonic unit uses a full-focusing method to perform pixel-by-pixel delay superposition of the full-matrix capture data to generate an internal crack image; the distributed optical fiber unit monitors the time-series data of concrete surface strain in real time and sets a strain gradient threshold; when the strain change exceeds the threshold, the distributed optical fiber unit triggers abnormal area marking and starts a high-precision detection mode.

[0091] The spatiotemporal alignment submodule for achieving time and spatial synchronization of multimodal data includes: the submodule is equipped with a GPS synchronization clock module to provide a unified timestamp for the SH array ultrasonic unit, distributed fiber optic unit, and millimeter-wave radar unit to achieve time synchronization; the submodule connects the distributed fiber optic demodulator and the millimeter-wave radar unit to an edge computing terminal via Ethernet to form a networked integration; the submodule uses a total station or built-in positioning module to record the physical coordinates of each unit to establish a unified three-dimensional coordinate system; and the submodule projects the scanning data of the millimeter-wave radar unit, the strain monitoring points of the distributed fiber optic unit, and the images of the SH array ultrasonic unit onto the same spatial grid to achieve spatial overlap.

[0092] The deep learning fusion submodule uses an improved Transformer architecture to perform cross-modal feature fusion to generate a 3D crack model from spatiotemporally aligned data. This includes: an improved Transformer architecture incorporating a spatiotemporal attention module; the spatial attention component of this module extracts local features from the internal crack images of the SH array ultrasonic unit and the surface crack images of the millimeter-wave radar unit through convolutional layers; the temporal attention component tracks the sequential changes in the strain time-series data of the distributed optical fiber unit to identify crack propagation trends; the deep learning fusion submodule employs a cross-modal attention mechanism to constrain different modal features within the same semantic space; and the deep learning fusion submodule generates a structured 3D crack model and its feature parameters based on the constrained features.

[0093] Regarding the environmental compensation and correction submodule's correction of environmental factors for the pre-fusion data, the following operations are involved: the environmental compensation and correction submodule collects temperature and humidity data in real time; the environmental compensation and correction submodule dynamically corrects the sound velocity of the SH array ultrasonic unit and the strain value of the distributed fiber optic unit based on the temperature and humidity data through Gaussian process regression; the environmental compensation and correction submodule uses an accelerometer to monitor the vibration frequency in real time; and the environmental compensation and correction submodule uses a band-stop filter to eliminate interference signals in the monitored vibration frequency band to improve the signal-to-noise ratio of the millimeter-wave radar unit.

[0094] In summary, this invention provides a concrete crack detection method based on multimodal sensing and deep learning. It integrates complementary data from three physical mechanisms: ultrasound (internal cracks), fiber optics (surface strain time series), and millimeter-wave radar (surface and shallow cracks). This overcomes the limitations of single sensors, such as susceptibility to interference and limited detection dimensions, significantly improving the accuracy and anti-interference capability of crack identification (especially internal and shallow cracks). Furthermore, it utilizes full-dimensional quantization to characterize multi-source data of cracks. It not only detects the presence of cracks but also generates a "digital crack profile" containing dimensional parameters (penetration, width), three-dimensional spatial coordinates, and predicted probabilities, achieving a transition from qualitative detection to quantitative and model-based diagnosis. Through environmental factor correction modules (temperature and humidity compensation, vibration compensation) and cross-modal attention mechanisms (noise suppression, feature consistency constraints), the influence of complex on-site environments on the data is effectively suppressed, ensuring the system's stability in real-world engineering scenarios and thus exhibiting stronger environmental robustness.

[0095] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, and all such modifications and additions should fall within the protection scope of the present invention.

Claims

1. A method for detecting concrete cracks based on multimodal sensing and deep learning, characterized in that, Includes the following steps: S10: Perform multimodal data acquisition and processing operations, and use the multimodal sensing module (10) composed of SH array ultrasonic unit (11), distributed optical fiber unit (12) and millimeter-wave radar unit (13) to detect the concrete structure; wherein, the SH array ultrasonic unit (11) acquires crack image data inside the concrete and records it as ultrasonic image data, the distributed optical fiber unit (12) acquires strain time series data on the concrete surface, and the millimeter-wave radar unit (13) acquires crack image data on the concrete surface and shallow layer and records it as radar image data; S20: Input the multi-source heterogeneous data collected in step S10 into the data fusion and correction module (20) for processing; this step specifically includes: S21: Provide a unified timestamp for the multimodal sensing module (10) by synchronizing the GPS clock, and establish a unified three-dimensional spatial coordinate system based on the pre-calibrated physical coordinates of each sensor, so as to map the crack image data, strain time series data and crack image data into the same spatiotemporal reference system; S22: Perform environmental factor correction on the data before fusion. Based on the temperature, humidity and vibration data collected by environmental sensors, perform environmental factor compensation correction on the multimodal data to reduce the impact of environmental changes on the crack characteristics. S23: An improved Transformer network is used to extract and fuse features from spatiotemporally aligned multimodal data to generate fused crack features; S30: Three-dimensional crack model generation step, based on the corrected crack features output in step S20, constructs and outputs a three-dimensional model of the internal cracks in the concrete and its feature parameters; the feature parameters include crack size parameters.

2. The method according to claim 1, characterized in that, The S10 step specifically includes: S101: Including the pre-scanning stage, firstly, the millimeter-wave radar unit (13) is used to quickly scan the concrete surface and mark the suspicious crack areas on the surface and in the shallow layer; S102: Includes the execution of a precise detection phase, controlling the SH array ultrasonic unit (11) to perform full-focus imaging on the marked suspicious area, while the distributed optical fiber unit (12) monitors the strain in real time; if the optical fiber strain data triggers a preset abnormal threshold, the high-precision detection mode of the SH array ultrasonic unit (11) is activated.

3. The method according to claim 1 or 2, characterized in that, In step S21, spatiotemporal alignment specifically includes: S211: The GPS synchronization clock module provides a unified UTC timestamp for the data acquisition actions of the SH array ultrasonic unit (11), distributed optical fiber unit (12) and millimeter-wave radar unit (13), so that the data of each sensor carries the same time reference. S212: Before detection, record the physical coordinates of each sensor and establish a unified three-dimensional coordinate system (X, Y, Z); transmit the scanning data coordinates of the millimeter-wave radar, the layout coordinates of the fiber optic strain monitoring points, and the geometric layout coordinates of the SH ultrasonic array to the edge computing terminal via Ethernet protocol and project them onto the same spatial grid.

4. The method according to claim 1, characterized in that, In step S23, the improved Transformer network utilizes a cross-modal attention mechanism to jointly model acoustic image features from the SH array ultrasonic unit (11), electromagnetic image features from the millimeter-wave radar unit (13), and strain time-series features from the distributed optical fiber unit (12) in the same semantic space.

5. A concrete internal crack detection system based on multimodal sensing and deep learning, used to implement the method described in any one of claims 1-4, characterized in that, include: The multimodal sensing module (10) includes an SH array ultrasonic unit (11), a distributed optical fiber unit (12), and a millimeter-wave radar unit (13), used to acquire multi-source heterogeneous detection data of concrete structures; the SH array ultrasonic unit (11) is used to acquire crack image data inside the concrete and record it as ultrasonic image data; the distributed optical fiber unit (12) is used to acquire strain time series data of the concrete surface; and the millimeter-wave radar unit (13) is used to acquire crack image data of the concrete surface and shallow layers and record it as radar image data. The data fusion correction module (20) is communicatively connected to the multimodal sensing module (10) and is used to process the multi-source heterogeneous detection data; The data fusion correction module (20) includes: The spatiotemporal alignment submodule (21) is used to provide a unified timestamp for the multimodal sensing module (10) through a GPS synchronization clock, and to establish a unified three-dimensional spatial coordinate system based on the pre-calibrated physical coordinates of each sensor, so as to map the ultrasonic image data, strain time series data and radar image data into the same spatiotemporal reference system. The environmental compensation and correction submodule (22) is used to perform environmental factor correction on the data before fusion. It uses temperature, humidity and vibration data collected by environmental sensors to correct the drift between ultrasonic image data and strain time series data through Gaussian process regression model, and eliminates the interference of vibration on radar image data through band-stop filter. Then it is re-fused to obtain the corrected crack features. The deep learning fusion submodule (23) is used to extract and fuse features of the spatiotemporally aligned multimodal data using an improved Transformer network. The improved Transformer network includes a spatial attention module to process ultrasonic image data and radar image data, and a temporal attention module to process fiber strain time series features, thereby generating fused crack features. The output module (30) is connected to the data fusion correction module (20) and is used to generate and output a three-dimensional model of the crack and its feature parameters based on the corrected crack features.

6. The system according to claim 5, characterized in that, The SH array ultrasound unit (11) includes: The shear horizontal wave transmitter is arranged in a ring, with an operating frequency of 100-500 kHz and an array element spacing of 10 cm, covering a detection grid of 0.5 m × 0.5 m; the signal processing unit is configured to acquire signals using full matrix acquisition technology and apply a full focusing method to convert the raw data into tomographic images.

7. The system according to claim 5, characterized in that, The distributed optical fiber unit (12) uses flexible packaging technology to directly attach single-mode optical fiber to the concrete surface. Its data processing unit is configured to trigger abnormal area marking by setting a strain gradient threshold in order to determine the location of cracks.

8. The system according to claim 5, characterized in that, The millimeter-wave radar unit (13) is a millimeter-wave synthetic aperture radar operating in the 60 GHz band. Its data processing unit is configured to perform a back-projection imaging algorithm on the acquired radar echo signal to generate a distribution image of the concrete surface and shallow cracks.

9. The system according to claim 5, characterized in that, The spatiotemporal alignment submodule (21) further includes a networked integration unit, which connects the demodulator of the distributed optical fiber unit (12), the millimeter-wave radar unit (13) and the edge computing terminal via Ethernet to form a high-speed data communication network for real-time transmission of sensor data with spatiotemporal coordinates.

10. The system according to claim 5, characterized in that, The environmental compensation correction submodule (23) includes: The temperature and humidity compensation unit is configured to use a Gaussian process regression model to dynamically correct the ultrasonic velocity and fiber strain values ​​based on real-time collected environmental temperature and humidity data. The vibration filtering unit is configured to use an accelerometer to monitor vibration and use a band-stop filter to eliminate interference frequency bands to radar signals.