Tunnel lining compactness defect detection method and system based on multi-modal data
Through multimodal data fusion and cross-modal comparison learning, a tunnel lining density defect detection system is built, which solves the problem of insufficient recognition accuracy under a single modal detection method, and achieves high accuracy and intelligent defect classification.
Patent Information
- Application Number
- CN202510632858.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The existing tunnel lining density defect detection methods mainly rely on single modal detection methods, with limited information dimensions, making it difficult to distinguish different types of density defects, and ignores the delay correlation and causal mechanism between multimodal data, limiting the recognition accuracy and classification capabilities.
Using a detection method based on multimodal data, a geological radar preliminary detection, combined with polarization response data, infrared active thermal imaging data and acoustic impact response data, feature extraction and encoding processing are carried out, a model-site-band ternary heterogeneous pattern is constructed, and a cross-modal comparison learning and sparse attention mechanism are combined, and defect categories are finally classified through diffusion mapping.
It significantly improves the comprehensiveness and accuracy of identification of compact defects, and can accurately distinguish concrete unfilled defects, filled stone defects and hole slag defects, which improves the reliability and intelligence of inspection.
Smart Images

Figure CN120180308A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tunnel lining detection. Specifically, it relates to a method and system for detecting the compactness defects of tunnel linings based on multi-modal data. Background Technique
[0002] For the quality inspection of tunnel linings, ground penetrating radar is mostly used as a preliminary diagnostic tool. This method penetrates the lining with high-frequency electromagnetic waves and analyzes their echo responses to identify whether there are compactness problems inside. However, ground penetrating radar can only determine whether there are compactness defects in the lining. For different types of defects (such as non-compact concrete, filled block stones and tunnel slag), it is difficult to effectively distinguish them due to the extremely similar electromagnetic wave responses. This makes it difficult to formulate targeted subsequent disease treatment strategies, affecting the practicality and accuracy of diagnosis.
[0003] Currently, the detection of lining defects mainly relies on single-modal detection means. The information dimensions obtained by it are limited, and there are significant ambiguities in the response characteristics of different types of compactness defects, resulting in insufficient classification accuracy. At the same time, existing research often ignores the potential time-delay correlation and causal mechanism between multi-modal data and lacks a systematic fusion modeling method, restricting the intelligent development of compactness defect identification.
[0004] Therefore, there is an urgent need for a new detection method that can fuse multi-modal perception data, fully explore the correlations between modalities, and improve the identification accuracy and classification ability of compactness defects. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for detecting the compactness defects of tunnel linings based on multi-modal data to improve the above problems. To achieve the above purpose, the technical solutions adopted by the present invention are as follows: In the first aspect, the present application provides a method for detecting the compactness defects of tunnel linings based on multi-modal data, including: Conduct preliminary quality inspection on the tunnel lining through ground penetrating radar to obtain the compactness defects to be detected; Obtain multi-modal data of the detected and to-be-detected compactness defects, where the multi-modal data includes polarization response data, infrared active thermographic data, and acoustic shock response data; Extract and encode the features of the multi-modal data to obtain a multi-modal embedded representation; Perform heterogeneous graph modeling and cross-modal contrast learning based on the multi-modal embedded representation to obtain a graph structure; Jointly process the graph structure based on the sparse attention mechanism and multi-modal fusion to obtain an embedded vector; Classify the embedded vectors based on diffusion mapping to obtain the classification results of the compactness defects to be detected, where the classification results include concrete non-compactness defects, filling blockstone defects, and hole slag defects.
[0006] In a second aspect, the present application also provides a tunnel lining compactness defect detection system based on multi-modal data, including: A preliminary detection unit for preliminarily detecting the quality of the tunnel lining through a ground penetrating radar to obtain the compactness defects to be detected; An acquisition unit for acquiring the multi-modal data of the detected and to-be-detected compactness defects, where the multi-modal data includes polarization response data, infrared active thermographic data, and acoustic shock response data; An encoding unit for performing feature extraction and encoding processing on the multi-modal data to obtain a multi-modal embedded representation; A modeling unit for performing heterogeneous graph modeling and cross-modal contrast learning based on the multi-modal embedded representation to obtain a graph structure; A joint processing unit for jointly processing the graph structure based on a sparse attention mechanism and multi-modal fusion to obtain an embedded vector; A classification unit for classifying the defect categories of the embedded vector based on diffusion mapping to obtain the classification results of the compactness defects to be detected, where the classification results include concrete non-compactness defects, filling blockstone defects, and hole slag defects.
[0007] The beneficial effects of the present invention are as follows: The present invention preliminarily detects the compactness defects of the tunnel lining through a ground penetrating radar, and combines multi-modal data to significantly improve the comprehensiveness and accuracy of compactness defect identification. To solve the scale and semantic differences of multi-modal data, a unified feature extraction and encoding strategy is proposed, and a three-way heterogeneous graph of modality - part - frequency band is constructed to effectively depict the spatial, frequency domain, and modal differences of defect responses. In addition, a cross-modal causal learning mechanism is introduced, and the time-delay correlation between different modalities is jointly modeled through mutual information and Granger causal analysis to enhance the graph modeling ability and the discriminability and interpretability of features. The defect response distribution modeling based on kernel density estimation can accurately depict the distribution characteristics of compactness defects, realize the adaptive setting of defect boundaries and discrimination thresholds, and further improve the reliability of compactness defect detection.
[0008] Other features and advantages of the present invention will be described in the subsequent specification, and some of them will become obvious from the specification, or can be understood by implementing the embodiments of the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings
[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0010] Figure 1 Schematic diagram of the process of the tunnel lining compactness defect detection method based on multi-modal data described in the embodiments of the present invention; Figure 2 Schematic diagram of the structure of the tunnel lining compactness defect detection system based on multi-modal data described in the embodiments of the present invention. Specific embodiments
[0011] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but only represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0012] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, terms such as "first" and "second" are only used for differential description and cannot be understood as indicating or implying relative importance.
[0013] Embodiment 1: This embodiment provides a tunnel lining compactness defect detection method based on multi-modal data.
[0014] See Figure 1 , which shows that this method includes step S1, step S2, step S3, step S4, step S5, and step S6.
[0015] Step S1: Conduct a preliminary quality inspection on the tunnel lining through a ground penetrating radar to obtain the compactness defects to be detected; In this embodiment, ground penetrating radar is used to conduct preliminary quality inspection on the tunnel lining to be detected, quickly obtain the areas with abnormal echo characteristics, and obtain typical quality defect results, such as crack defects, cavities or voids, deformations, and compactness defects. Since the abnormal echo characteristics of different quality defect results are different, the compactness defect of the tunnel lining to be detected and the position information of the compactness defect to be detected can be obtained.
[0016] Step S2: Obtain multi-modal data of the detected and to-be-detected compactness defects, where the multi-modal data includes polarization response data, infrared active thermography data, and acoustic shock response data; In this embodiment, the multi-modal data of the detected compactness defects provides label references and feature priors for subsequent contrast learning, graph structure construction, and classification, etc. And the samples of the detected compactness defects can be obtained from the defect database of existing engineering projects.
[0017] Since a single modality has limitations and cannot fully distinguish the types of compactness defects, using multi-modal data can make up for their respective weaknesses and improve the ability to distinguish compactness defects. For example, the heat conduction characteristics, polarization response characteristics, and acoustic wave propagation characteristics of concrete non-compactness defects, filled block stone defects, and cavity slag defects are not exactly the same.
[0018] Step S3: Perform feature extraction and encoding processing on the multi-modal data to obtain a multi-modal embedding representation; In this embodiment, since the dimensions, scales, and distributions of the multi-modal data are different, directly splicing or comparing these data is likely to introduce inter-modal inconsistencies, resulting in chaotic features and unstable training. Therefore, it is necessary to perform feature extraction and encoding processing to extract representative and dimensionally consistent feature vectors for each, reducing redundancy. Convert the original multi-modal data with different sources and physical meanings into a representation form that can be uniformly analyzed and fused, that is, a multi-modal embedding representation.
[0019] In step S3, the obtaining of the multi-modal embedding representation includes: Step S31: Conduct polarization feature analysis on the polarization response data to obtain a polarization response vector; In this embodiment, the echo amplitudes in different polarization states in the polarization response data are obtained, a polarization scattering matrix is constructed through the echo amplitudes, and then the polarization entropy, homogeneity, and main polarization direction angle are calculated through the polarization scattering matrix. A polarization response vector is constructed through the polarization entropy, homogeneity, and main polarization direction angle.
[0020] Step S32: Construct a temperature gradient tensor based on the infrared active thermography data, and extract a thermal anomaly response vector through the temperature gradient tensor; In this embodiment, the infrared active thermographic data is a sequence of multiple infrared thermographic images. The temperature gradient tensor is calculated from the infrared active thermographic data in the time series. A heat diffusion model is established through the heat conduction equation, and then the thermal conductivity and thermal inertia are inversely deduced from the temperature gradient tensor.
[0021] Based on the thermal conductivity and thermal inertia, a thermal conductivity map and a thermal inertia distribution map are constructed. Local anomalies (such as mutations and gradient extremes) in the thermal conductivity map and the thermal inertia distribution map are used to extract the thermal anomaly response feature items (thermal inertia distribution deviation value, thermal conductivity gradient mutation index, local heat diffusion coefficient estimation value, and heat flow direction offset angle statistics), and a thermal anomaly response vector is obtained. Therefore, the thermal anomaly response vector can reflect the thermal property anomaly characteristics inside the tunnel lining caused by factors such as decreased compactness, foreign object embedding, or material replacement.
[0022] Step S33: Construct an impact echo time-delay spectrum based on the acoustic wave impact response data, and extract an acoustic wave impact response vector from the impact echo time-delay spectrum. The acoustic wave impact response vector includes a reflection amplitude statistical matrix and an energy distribution texture feature. In this embodiment, the acoustic wave impact response data is a time-domain signal representing the acoustic wave signal intensity at different time points. The original acoustic wave impact response data is preprocessed to obtain an impact signal and an echo signal. The preprocessing includes denoising, filtering, and removing unnecessary background noise. Using the cross-correlation analysis or the matched filtering method, the impact signal and the echo signal are subjected to a correlation analysis to obtain the impact echo time-delay spectrum.
[0023] The reflection amplitudes at different time points are extracted from the impact echo time-delay spectrum, and a statistical matrix, that is, the reflection amplitude statistical matrix, is constructed to describe the statistical characteristics of the reflection intensity of the acoustic wave echo signal at different time delays. By calculating the local energy distribution of the impact echo time-delay spectrum, the energy distribution texture feature is obtained to describe the energy distribution characteristics of the acoustic wave echo on the time axis. Finally, combining the reflection amplitude statistical matrix and the energy distribution texture feature, a complete acoustic wave impact response vector is constructed.
[0024] Step S34: Encode the polarization response vector, the thermal anomaly response vector, and the acoustic wave impact response vector through a multi-modal autoencoder to obtain a multi-modal embedding representation.
[0025] In this embodiment, the variational autoencoder is selected as the multi-modal autoencoder. For the response vectors of each modality (polarization response vector, thermal anomaly response vector, and acoustic wave impact response vector), independent variational autoencoders are used for encoding to form a latent space distribution, that is, the embedding representation corresponding to the modality is obtained. Finally, the embedding representations of all modalities are combined into a multi-modal embedding representation. Among them, the modalities include the polarization modality, the thermographic modality, and the acoustic wave modality.
[0026] Step S4: Perform heterogeneous graph modeling and cross-modal contrastive learning based on the multi-modal embedding representation to obtain a graph structure; In this embodiment, since the responses of the same defect in different modalities have information sharing and complementarity, a graph structure is used and a cross-modal contrastive learning mechanism is introduced to improve the structural rationality and semantic discrimination of feature fusion between different modalities, and this cross-source dependence relationship can be more realistically modeled.
[0027] In step S4, the performing heterogeneous graph modeling and cross-modal contrastive learning based on the multi-modal embedding representation to obtain a graph structure includes: Step S41: Construct a ternary heterogeneous graph for each density defect with respect to modality - part - frequency band based on the multi-modal embedding representation; In this embodiment, corresponding graph nodes are constructed according to each density defect. Each graph node is a ternary heterogeneous graph, that is, it includes modality information, part information, and frequency band information. The modality information is the multi-modal embedding representation, including the embedding representation of each modality. Among them, the perception mechanisms of each modality are different, and the sensitivities to different types of density defects are also different. There are differences in the material density of different parts (the crown, side walls, and shoulders) of the tunnel lining, and the damage mechanisms are also different. The frequency band controls the information penetration depth and resolution, supplementing the deficiencies of the modality itself. For example, the same type of density defect will have different manifestations in the modal response due to factors such as structural differences, stress differences, and material property differences at different parts (such as the crown, side walls, shoulders, etc.). Through the distinction of part information, this spatial difference can be better understood and captured, so as to more accurately identify and classify defect types.
[0028] Therefore, through the ternary heterogeneous graph, these three types of information are explicitly incorporated into the graph modeling, which not only increases the integrity of the structural expression, but also improves the accuracy and interpretability of cross-modal defect recognition. And the part information can perform a structural semantic mapping through the coordinates of the density defect, mapping the coordinates to the corresponding structural part labels.
[0029] Step S42: Extract the time-delay correlation between different modalities based on an improved multi-modal causal discovery algorithm, and generate a weight matrix for the ternary heterogeneous graph based on the time-delay correlation; In this embodiment, the weight matrix describes the relationship between the modal responses inside each graph node, which not only expresses the internal characteristics of the node itself, but also reflects the multi-modal characteristics of the node in subsequent calculations.
[0030] In step S42, the generation steps of the weight matrix are as follows: Step S421: Align the polarization response vector, thermal anomaly response vector, and acoustic shock response vector to obtain aligned vectors of multiple modalities; Step S422: Perform sliding window segmentation on the aligned vectors to obtain continuous vector segments; Step S423: Calculate the mutual information between different modalities based on consecutive vector segments; In this embodiment, the mutual information between consecutive vector segments is calculated to measure the amount of information shared by different modalities within the same time period. The greater the mutual information, the stronger the correlation between the two modalities. Assume modality and modality The vector segments on window are and , then the calculation formula for mutual information is: ; ; In the formula, represents and 's mutual information, and respectively represent and 's information entropy, represents and 's joint entropy, represents the vector segment of modality on window , represents the vector segment of modality on window , represents the mutual information between modality and modality , represents the total number of windows.
[0031] Step S424: Conduct Granger causality analysis based on consecutive vector segments to obtain the causality strength between different modalities; In this embodiment, for each window, two sets of models are established, namely the non-causal model and the causal model. Specifically: ; ; In the formula, represents the value of the first vector segment of modality at time on window , represents the value of the second vector segment of modality at time on window , represents the regression order, that is, how many time steps of information are used for modeling, Indicating modality At the moment on the window At the moment The time-delayed moment Of the first vector segment value Indicating modality At the moment on the window At the moment The time-delayed moment Of the second vector segment value Indicating modality At the moment on the window At the moment The time-delayed moment Of the second vector segment value Indicating modality Of the autoregressive coefficient Indicating modality For the modality Of the causal influence coefficient Indicating modality Of the residuals of the non-causal model Indicating the existence of modality The modality with influence Of the residuals of the causal model
[0032] Where there is a causal model with the existence of modality For the modality Of the influence. Therefore, let the residual variances of the non-causal model and the causal model be And respectively. Then the causal strength is obtained as: ; ; In the formula, Indicates the causal strength of modality On the window For modality To modality Indicates the causal strength of modality To modality And Indicates the total number of windows
[0033] Step S425: Construct the time-delay correlation between different modalities through mutual information and causal strength, and construct the weight matrix of the tripartite heterogeneous graph with the time-delay correlation as the edge weight
[0034] In this embodiment, the formula for the time-delay correlation between two modalities is: ; ; In the formula, Represents a modality and the modality of the time delay correlation Represents a modality and the modality of the mutual information and both represent weight parameters Represents a modality and the modality of the causality strength Represents a modality with respect to the modality of the causality strength Represents a modality with respect to the modality of the causality strength
[0035] Step S43: Based on the physical properties and structural design of the tunnel lining, perform physical constraint pruning and correction on the ternary heterogeneous graph to obtain a multi-modal response graph; In this embodiment, through the physical properties and structural design of the tunnel lining, unreasonable or redundant connection relationships in the ternary heterogeneous graph can be eliminated. For example, edges with mutual information or causality strength less than a threshold are regarded as invalid.
[0036] Step S44: Optimize the representation consistency and structural contrast ability in the multi-modal response graph through cross-modal contrast learning, and construct a graph structure with each compactness defect as a graph node and the similarity between defects as the edge weight.
[0037] In this embodiment, each graph node includes not only its embedded representations in multiple modalities, but also a weight matrix representing the relationship between the internal modal responses of the graph node. Calculate the embedding similarity of two graph nodes through the embedded representation and the weight matrix, calculate the spatial distance attenuation term and the frequency band similarity adjustment term through the part information and the frequency band information, and then construct the edge weight between two graph nodes, that is, the similarity between defects, through the embedding similarity, the spatial distance attenuation term, and the frequency band similarity adjustment term.
[0038] Based on whether the nodes come from the same true defect category, construct positive and negative sample pairs, and optimize the multi-modal embedded representation through a contrast learning loss function (such as the InfoNCE loss function), so that the edge weight between positive sample node pairs is enhanced and the edge weight of negative sample pairs is weakened, thereby improving the representation consistency and structural distinguishability in the multi-modal response graph, and finally obtaining a graph structure with each compactness defect as a graph node and the similarity between defects as the edge weight. Therefore, the node feature vector of each graph node in the graph structure is composed of multi-modal embedded representation, weight matrix, part information, and frequency band information.
[0039] Step S5: Jointly process the graph structure based on the sparse attention mechanism and multimodal fusion to obtain an embedding vector; In step S5, the obtaining of the embedding vector includes: Step S51: Normalize the graph structure to generate a sparse graph representation. The normalization process includes node feature normalization, edge weight matrix sparsification, and structural adjacency matrix standardization; In this embodiment, the node feature vectors of each graph node in the graph structure are normalized to obtain normalized node feature vectors. When sparsifying the edge weight matrix, the edges with edge weights lower than the preset threshold between graph nodes are set to zero, and only the strongly similar connections are retained to form a sparse adjacency matrix. Calculate the degree matrix corresponding to the adjacency matrix, and obtain the normalized sparse adjacency matrix in a symmetric normalization form. Update the graph structure through the normalized node feature vectors and the normalized sparse adjacency matrix to obtain a sparse graph representation.
[0040] Step S52: Construct modal channels for each modality based on the sparse graph representation; In this embodiment, the normalized node feature vectors of each graph node in the sparse graph representation are split to obtain feature sub-vectors for each modality. For each modality, the feature sub-vectors of the corresponding modality of all graph nodes are extracted to form the feature matrix of the modal channel. And each modal channel shares the same normalized sparse adjacency matrix, indicating the topological invariance of the sparse graph representation.
[0041] Step S53: Dynamically calculate and aggregate the neighbor information of each graph node in the sparse graph representation based on the sparse attention mechanism; In this embodiment, first calculate the attention coefficients between each graph node and its corresponding neighbor nodes. Specifically, the attention coefficients are calculated using the attention mechanism through the normalized node feature vectors of the graph node and its corresponding neighbor nodes and the edge weights (determined by the normalized sparse adjacency matrix).
[0042] After obtaining the attention coefficients between the graph node and each of its corresponding neighbor nodes, the normalized node feature vectors of each neighbor node are weighted and aggregated through an activation function to obtain the neighbor information of each graph node.
[0043] Step S54: For each modal channel, obtain the embedding feature representation of each graph node through the neighbor information of each graph node; In this embodiment, the feature matrix of each modal channel is weighted and aggregated and updated through the neighbor information of each graph node to obtain the embedding feature representation of each graph node on each modal channel. This process is actually to update the features of the nodes through the neighbor information.
[0044] Step S55: Dynamically allocate the weighted coefficients for each modal channel based on the global modal attention mechanism, and fuse the embedded feature representations of each modal channel based on the weighted coefficients to obtain the embedded vector of each graph node.
[0045] In this embodiment, when using the global modal attention mechanism, for each modal channel, calculate the weighted coefficient according to the embedded feature representation of each graph node, and then perform weighted fusion on the embedded feature representations of each modal channel to obtain the embedded vector of each graph node. Therefore, this embedded vector fuses the embedded features of all modal channels and assigns the importance weights of each modality through the global modal attention mechanism, realizing the weighted integration of multimodal information.
[0046] Step S6: Classify the embedded vectors based on diffusion mapping to obtain the classification results of the density defects to be detected, and the classification results include concrete non-dense defects, filling block stone defects, and hole slag defects.
[0047] In step S6, obtaining the classification results of the density defects to be detected includes: Step S61: Construct a diffusion operator based on the embedded vector, and perform eigen-decomposition based on the diffusion operator to map each embedded vector to a manifold space point in the low-dimensional manifold space, and the manifold space points include detected and to-be-detected manifold space points; In this embodiment, based on the embedded vector of each graph node, use the Gaussian kernel function to construct the similarity matrix between graph nodes, and then construct the degree matrix through the similarity matrix, and obtain the diffusion operator based on the Markov diffusion process. Specifically: ; In the formula, represents the diffusion operator, represents the degree matrix, represents the similarity matrix.
[0048] Perform eigen-decomposition on the diffusion operator to obtain multiple eigenvectors and corresponding eigenvalues, constituting the manifold space points in the low-dimensional manifold space. Specifically: ; In the formula, represents the coordinate of the graph node in the low-dimensional manifold space under the hyperparameter , represents the th eigenvalue, represents the th component of the eigenvector corresponding to the th eigenvalue, where the hyperparameter is used to control the intensity of diffusion.
[0049] Step S62: Use kernel density estimation and the detected manifold space points to construct the probability density function of each defect category on the low-dimensional manifold space; In this embodiment, for each defect category, the detected manifold space points (i.e., the manifold space points corresponding to the detected density defects) are used to estimate the probability density function of this defect category on the low-dimensional manifold space. In this step, the Gaussian kernel function is used for estimation, and the specific formula is: ; In the formula, represents the probability density function of the defect category , represents the number of detected manifold space points of the defect category , represents the bandwidth parameter, which is used to control the width of the kernel function, represents the dimension of the low-dimensional manifold space, represents the defect category 's th coordinate of the detected manifold space point, represents the coordinate of the manifold space point to be detected, represents the Euclidean distance.
[0050] Step S63: Calculate the probability density of the manifold space point to be detected on each defect category based on the probability density function, and perform normalization processing on the probability density to obtain the normalized category probability; In this embodiment, the probability density of the manifold space point to be detected on each defect category is calculated through the probability density function. In order to make the sum of the probability densities of all defect categories equal to 1, the probability densities of each defect category of each point to be detected are normalized to obtain the normalized category probability of each defect category.
[0051] Step S64: Select the defect category with the highest normalized category probability as the defect category of the manifold space point to be detected, and obtain the classification result of the density defect to be detected.
[0052] In summary, the present invention conducts preliminary detection on the tunnel lining through ground penetrating radar, quickly locates potential density defect areas, and then combines multi-modal sensing technologies such as polarization response, infrared active thermography, and acoustic shock to obtain multi-source response data of the lining, significantly improving the comprehensiveness and accuracy of defect identification. Aiming at the scale and semantic differences between multi-modal data, a unified feature extraction and encoding strategy is proposed to obtain multi-modal embedded representations, and a modal-part-frequency three-way heterogeneous graph is constructed to effectively characterize the differences of defect responses in space, frequency domain, and modality.
[0053] In addition, the present invention further introduces a cross-modal causal learning mechanism, jointly models the time-delay correlation between different modalities through mutual information and Granger causality analysis, constructs a graph weight matrix with reasonable structure and accurate weights, enhances the discriminability and interpretability of features while improving the graph modeling expression ability. Through the heterogeneous graph modeling and cross-modal contrast learning mechanism, not only can complementary information of each modality be fused, but also potential temporal causal structures can be fully mined. By introducing the defect response distribution modeling based on kernel density estimation, the probability density distribution characteristics of different types of compactness defects in the embedding space can be accurately characterized, realizing the boundary modeling of the defect area and the adaptive setting of the discrimination threshold, and improving the refinement level of defect classification.
[0054] Embodiment 2: As Figure 2 shown, this embodiment provides a tunnel lining compactness defect detection system based on multi-modal data. The system includes: A preliminary detection unit for performing preliminary quality detection on the tunnel lining through a ground penetrating radar to obtain the compactness defects to be detected; An acquisition unit for acquiring multi-modal data of the detected and to-be-detected compactness defects, where the multi-modal data includes polarization response data, infrared active thermography data, and acoustic shock response data; An encoding unit for performing feature extraction and encoding processing on the multi-modal data to obtain a multi-modal embedding representation; A modeling unit for performing heterogeneous graph modeling and cross-modal contrast learning based on the multi-modal embedding representation to obtain a graph structure; A joint processing unit for jointly processing the graph structure based on a sparse attention mechanism and multi-modal fusion to obtain an embedding vector; A classification unit for classifying the defect categories of the embedding vector based on diffusion mapping to obtain the classification result of the to-be-detected compactness defects, where the classification result includes concrete non-compactness defects, filling blockstone defects, and tunnel slag defects.
[0055] The modeling unit includes: A first construction subunit for constructing a ternary heterogeneous graph of each compactness defect with respect to modality - part - frequency band based on the multi-modal embedding representation; An extraction subunit for extracting the time-delay correlation between different modalities based on an improved multi-modal causal discovery algorithm and generating a weight matrix of the ternary heterogeneous graph based on the time-delay correlation; A correction subunit for performing physical constraint pruning and correction on the ternary heterogeneous graph based on the physical properties and structural design of the tunnel lining to obtain a multi-modal response graph; The second construction subunit is used to optimize the representation consistency and structural contrast ability in the multi-modal response map through cross-modal contrast learning, and construct a graph structure with each density defect as a graph node and the similarity between defects as edge weights.
[0056] The joint processing unit includes: A processing subunit is used to perform normalization processing on the graph structure to generate a sparse graph representation. The normalization processing includes node feature normalization, edge weight matrix sparsification, and structural adjacency matrix standardization; The third construction subunit is used to construct the modal channels of each modality based on the sparse graph representation; The first calculation subunit is used to dynamically calculate and aggregate the neighbor information of each graph node in the sparse graph representation based on the sparse attention mechanism; The second calculation subunit is used to obtain the embedded feature representation of each graph node through the neighbor information of each graph node for each modal channel; The third calculation subunit is used to dynamically assign the weighted coefficients of each modal channel based on the global modal attention mechanism, and fuse the embedded feature representations of each modal channel based on the weighted coefficients to obtain the embedded vector of each graph node.
[0057] The classification unit includes: A decomposition subunit is used to construct a diffusion operator based on the embedded vector, and perform eigenvalue decomposition based on the diffusion operator to map each embedded vector to a manifold space point in the low-dimensional manifold space. The manifold space points include detected and undetected manifold space points; The fourth calculation subunit is used to construct the probability density function of each defect category in the low-dimensional manifold space using kernel density estimation and the detected manifold space points; The fifth calculation subunit is used to calculate the probability density of the undetected manifold space points on each defect category based on the probability density function, and perform normalization processing on the probability density to obtain the normalized category probability; The classification subunit is used to select the defect category with the highest normalized category probability as the defect category of the undetected manifold space point to obtain the classification result of the undetected density defect.
[0058] It should be noted that regarding the system in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0059] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0060] As described above, it is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A tunnel lining compactness defect detection method based on multimodal data, characterized in that: include: Conduct preliminary quality inspection of tunnel lining by geological radar to obtain the compactness defects to be detected; Acquiring multimodal data of detected and to-be-detected compactness defects, wherein the multimodal data includes polarization response data, infrared active thermal imaging data, and acoustic wave impact response data; Perform feature extraction and encoding on multimodal data to obtain multimodal embedding representation; Based on multimodal embedding representation, heterogeneous graph modeling and cross-modal comparative learning are performed to obtain the graph structure; The graph structure is jointly processed based on sparse attention mechanism and multimodal fusion to obtain the embedding vector; The embedded vectors are classified into defect categories based on diffusion mapping to obtain classification results of compactness defects to be detected, wherein the classification results include concrete non-compactness defects, filling block stone defects and cavity slag defects.
2. The tunnel lining compactness defect detection method based on multimodal data according to claim 1 is characterized in that: The multimodal embedding representation is obtained, including: Perform polarization characteristic analysis on the polarization response data to obtain a polarization response vector; Construct a temperature gradient tensor based on infrared active thermal imaging data, and extract thermal anomaly response vectors through the temperature gradient tensor; Constructing an impact echo delay spectrum based on the acoustic wave impact response data, and extracting an acoustic wave impact response vector through the impact echo delay spectrum, wherein the acoustic wave impact response vector includes a reflection amplitude statistical matrix and an energy distribution texture feature; The polarization response vector, thermal anomaly response vector and acoustic shock response vector are encoded through a multimodal autoencoder to obtain a multimodal embedding representation.
3. The tunnel lining compactness defect detection method based on multimodal data according to claim 2 is characterized in that: The heterogeneous graph modeling and cross-modal comparative learning based on multimodal embedding representation are performed to obtain a graph structure, including: Based on the multimodal embedding representation, a ternary heterogeneous graph of each density defect with respect to mode, location and frequency band is constructed; Based on the improved multimodal causal discovery algorithm, the time delay correlation between different modes is extracted, and the weight matrix of the ternary heterogeneous graph is generated based on the time delay correlation; Based on the physical properties and structural design of the tunnel lining, the ternary heterogeneous graph is physically pruned and modified to obtain a multimodal response graph. Through cross-modal contrastive learning, the representation consistency and structural contrast ability in multimodal response graphs are optimized, and a graph structure is constructed with each density defect as a graph node and the similarity between defects as the edge weight.
4. The tunnel lining compactness defect detection method based on multimodal data according to claim 3 is characterized in that: The steps for generating the weight matrix are: Aligning the polarization response vector, the thermal anomaly response vector and the acoustic wave impact response vector to obtain alignment vectors of multiple modes; Perform sliding window segmentation on the aligned vector to obtain continuous vector segments; Compute mutual information between different modalities based on consecutive vector segments; Granger causality analysis is performed based on continuous vector segments to obtain the causal strength between different modes; The time delay correlation between different modes is constructed through mutual information and causal strength, and the weight matrix of the ternary heterogeneous graph is constructed using the time delay correlation as the edge weight.
5. The tunnel lining compactness defect detection method based on multimodal data according to claim 1 is characterized in that: The step of obtaining the embedding vector comprises: Normalizing the graph structure to generate a sparse graph representation, wherein the normalization includes node feature normalization, edge weight matrix sparseness, and structural adjacency matrix standardization; Construct modal channels for each modality based on sparse graph representation; Dynamically calculate and aggregate the neighbor information of each graph node in the sparse graph representation based on the sparse attention mechanism; For each modal channel, the embedded feature representation of each graph node is obtained through the neighbor information of each graph node; Based on the global modal attention mechanism, the weight coefficient of each modal channel is dynamically allocated, and the embedded feature representation of each modal channel is fused based on the weight coefficient to obtain the embedding vector of each graph node.
6. The tunnel lining compactness defect detection method based on multimodal data according to claim 1 is characterized in that: The obtaining of the classification result of the compactness defect to be detected includes: A diffusion operator is constructed based on the embedded vector, and eigendecomposition is performed based on the diffusion operator to map each embedded vector to a manifold space point in a low-dimensional manifold space, wherein the manifold space point includes a detected manifold space point and a manifold space point to be detected; The probability density function of each defect category in the low-dimensional manifold space is constructed using kernel density estimation and the detected manifold space points; The probability density of the manifold space point to be detected on each defect category is calculated based on the probability density function, and the probability density is normalized to obtain the normalized category probability; The defect category with the highest normalized category probability is selected as the defect category of the manifold space point to be detected, and the classification result of the density defect to be detected is obtained.
7. A tunnel lining compactness defect detection system based on multimodal data, characterized in that: include: The preliminary inspection unit is used to conduct preliminary quality inspection of the tunnel lining by using geological radar to obtain the compactness defects to be detected; An acquisition unit, used to acquire multimodal data of the compactness defects that have been detected and to be detected, wherein the multimodal data includes polarization response data, infrared active thermal imaging data, and acoustic impact response data; The encoding unit is used to extract features and encode multimodal data to obtain a multimodal embedding representation; A modeling unit, which is used to perform heterogeneous graph modeling and cross-modal comparative learning based on multimodal embedding representation to obtain the graph structure; A joint processing unit, which is used to jointly process the graph structure based on the sparse attention mechanism and multimodal fusion to obtain an embedding vector; The classification unit is used to classify the embedded vector into defect categories based on diffusion mapping to obtain the classification results of the compactness defects to be detected, wherein the classification results include concrete non-compactness defects, filling block stone defects and cavity slag defects.
8. The tunnel lining compactness defect detection system based on multimodal data according to claim 7 is characterized in that: The modeling unit comprises: A first construction subunit is used to construct a ternary heterogeneous graph of each density defect with respect to mode-position-frequency band based on a multimodal embedding representation; An extraction subunit, used to extract the time delay correlation between different modes based on the improved multimodal causal discovery algorithm, and generate a weight matrix of the ternary heterogeneous graph based on the time delay correlation; The correction subunit is used to perform physical constraint pruning and correction on the ternary heterogeneous graph based on the physical properties and structural design of the tunnel lining to obtain a multimodal response graph; The second construction subunit is used to optimize the representation consistency and structural contrast ability in the multimodal response graph through cross-modal contrast learning, and to construct a graph structure with each density defect as a graph node and the similarity between defects as the edge weight.
9. The tunnel lining compactness defect detection system based on multimodal data according to claim 7 is characterized in that: The joint processing unit comprises: A processing subunit, used for normalizing the graph structure to generate a sparse graph representation, wherein the normalization includes node feature normalization, edge weight matrix sparseness, and structural adjacency matrix standardization; A third construction subunit is used to construct a modality channel of each modality based on the sparse graph representation; A first computing subunit, for dynamically computing and aggregating neighbor information of each graph node in the sparse graph representation based on a sparse attention mechanism; A second computing subunit is used to obtain, for each modal channel, an embedded feature representation of each graph node through neighbor information of each graph node; The third computing subunit is used to dynamically allocate the weight coefficient of each modal channel based on the global modal attention mechanism, and fuse the embedded feature representation of each modal channel based on the weight coefficient to obtain the embedding vector of each graph node.
10. The tunnel lining compactness defect detection system based on multimodal data according to claim 7, characterized in that: The taxonomic units include: A decomposition subunit, configured to construct a diffusion operator based on the embedded vector, and perform feature decomposition based on the diffusion operator to map each embedded vector to a manifold space point in a low-dimensional manifold space, wherein the manifold space point includes a detected manifold space point and a manifold space point to be detected; A fourth computing subunit, for constructing a probability density function of each defect category in the low-dimensional manifold space using kernel density estimation and detected manifold space points; A fifth calculation subunit is used to calculate the probability density of the manifold space point to be detected on each defect category based on the probability density function, and normalize the probability density to obtain a normalized category probability; The classification subunit is used to select the defect category with the highest normalized category probability as the defect category of the manifold space point to be detected, and obtain the classification result of the density defect to be detected.
Citation Information
Patent Citations
Method, device, equipment and medium for improving geological radar detection precision of tunnel lining defects
CN117538863A
Tunnel lining damage detection method, system and equipment based on data fusion and medium
CN119470438A
Tunnel lining defect rapid identification model based on geological radar method
CN211506935U
Tunnel lining vault void monitoring device and detection method
WO2024250421A1
Cited By
Tunnel lining leakage infrared-millimeter wave fusion detection method and system
CN120491045A