Composite material damage terahertz imaging method based on two-stage unsupervised learning
By employing a two-stage unsupervised learning framework and physics-driven signal preprocessing, the problems of data dependence and model generalization in composite material damage detection are solved, achieving efficient and accurate damage identification and imaging, which is suitable for industrial inspection of composite materials.
Patent Information
- Application Number
- CN202511273941.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-12-02
AI Technical Summary
Existing technologies for composite material damage detection rely on a large amount of manually labeled data, resulting in high computational complexity and insufficient model generalization ability. This makes it difficult to effectively identify damage under different test conditions, and the imaging resolution and accuracy are low.
A two-stage unsupervised learning approach is adopted, which generates a pseudo-label dataset through stacked autoencoders and spectral clustering, uses convolutional neural networks for unlabeled terahertz signal classification, and combines physical priors for signal alignment to achieve damage feature extraction and high-resolution imaging.
Training datasets can be built without manual annotation, which improves the model's generalization ability and imaging accuracy, reduces dependence on data volume, adapts to different testing conditions, and quickly responds to the needs of industrial sites.
Smart Images

Figure CN121049201A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of nondestructive testing of composite materials, and more specifically, to a terahertz imaging method for composite material damage based on two-stage unsupervised learning. Background Technology
[0002] Composite materials, due to their superior properties such as lightweight, corrosion resistance, high strength, and high stiffness, have become indispensable materials in high-performance applications such as aerospace, automotive, construction, and military. However, composite materials are susceptible to various types of damage caused by complex service environments and mechanical loads, such as invisible impact damage, matrix cracks, voids, delamination, fiber breakage, and misalignment. These damages can compromise structural integrity, and if not detected and repaired in time, may lead to performance degradation or even catastrophic failure. Therefore, timely detection and assessment of internal damage in composite materials are crucial for ensuring structural integrity and service safety.
[0003] Non-destructive testing (NDT) techniques are widely used for the detection and assessment of internal damage in composite material structures, such as ultrasonic testing, eddy current testing, infrared testing, and X-ray testing. However, these methods each have their limitations. For example, ultrasonic testing requires a coupling agent, eddy current testing is suitable for conductive materials, infrared testing is limited by ambient temperature, and X-ray testing is restricted by radiation issues. With the development of high-end equipment in aerospace, nuclear power, and other fields, higher demands are placed on NDT techniques. Therefore, there is an urgent need for a reliable, effective, and high-resolution NDT method to achieve accurate identification of internal material damage. Terahertz NDT, with a frequency range from 100 GHz to 10 THz, has attracted widespread attention in the field of composite material damage detection due to its high resolution, high penetration, reliability, and safety. During terahertz testing, when a terahertz pulse passes through or is reflected from the target sample, its amplitude, phase, and spectral characteristics are modulated by the physicochemical properties of the material. Therefore, the terahertz echo signal carries rich structural and damage information about the material's interior. When there is damage inside the material under test, it will cause changes in the characteristics of the terahertz echo signal. Based on these changes in characteristics, the damage can be located and imaged.
[0004] Currently, although terahertz nondestructive testing technology has received widespread attention, accurately and efficiently extracting damage features from complex terahertz response signals still faces some challenges, especially for complex multilayer composite structures. In fact, during terahertz testing, the terahertz response signal is composed of multiple echo components from different reflecting interfaces, including the damage interface and interlayer interfaces, leading to aliasing between the damage echo components and the interlayer interface echo components, increasing the difficulty of damage feature extraction. Furthermore, terahertz waves are susceptible to complex interferences such as dispersion, multiple reflections, and noise, inevitably causing overlap and distortion in the terahertz response signal, further increasing the difficulty of damage feature extraction. Currently, traditional terahertz signal processing methods, including wavelet transform, deconvolution, autoregressive spectrum estimation, and sparse representation, have been successfully used to address these problems. However, these methods are highly dependent on expert experience and manual intervention, and their performance is limited by the selection of optimal hyperparameters, making them difficult to handle complex terahertz signals. In addition, these methods experience increased computational complexity and significantly reduced computational efficiency when processing batches of terahertz data.
[0005] In recent years, deep learning, as an end-to-end feature extraction technology, has been applied to various fields due to its powerful feature learning capabilities. Currently, some research attempts to combine deep learning with terahertz nondestructive testing (NDT) to improve the efficiency and accuracy of internal damage detection in composite materials. However, it is worth noting that establishing a general and efficient deep learning framework for identifying composite material damage in practical applications of terahertz NDT still faces significant challenges. In fact, current deep learning models still need to address two thorny issues: model generalization and the need for a large amount of labeled terahertz test data. Specifically, traditional deep learning models can capture identifiable damage features by training on existing terahertz data from a limited set of test samples. However, during terahertz testing, the terahertz signal changes due to various factors, including the sample's material, structure, surface uniformity, damage thickness and depth, as well as test conditions and parameter settings. This leads to significant differences in data distribution between existing and changing terahertz data, even for the same sample. In this situation, since the changing terahertz data can be considered unseen data, the generalization ability of traditional models decreases, making it difficult to effectively identify damage. Furthermore, in terahertz nondestructive testing, current traditional deep learning models rely excessively on supervised learning. Although supervised learning is widely used in deep learning due to its superior feature learning capabilities, its performance is limited by large amounts of labeled data.
[0006] Therefore, this invention aims to propose a general and efficient unsupervised deep learning framework that does not rely on any human intervention or data labels, to achieve intelligent identification and high-resolution imaging of internal damage in composite materials under different application scenarios, thereby promoting the engineering application of deep learning technology in the field of practical terahertz nondestructive testing. Summary of the Invention
[0007] The purpose of this invention is to solve the problems mentioned in the background art, and therefore proposes a terahertz imaging method for composite material damage based on two-stage unsupervised learning.
[0008] The technical solution adopted by this invention to solve its technical problem is: A terahertz imaging method for composite material damage based on two-stage unsupervised learning is characterized by the following steps: S1. Prepare composite laminate samples with different damage thicknesses, obtain time-domain terahertz signals of different samples based on terahertz time-domain spectroscopy system, and construct multiple unlabeled single-sample damage terahertz datasets. S2. Construct a data alignment layer based on the prior of the terahertz signal to perform data alignment processing on the acquired terahertz time-domain signal; S3. Construct a two-stage unsupervised learning framework, including clustering and classification processes. Use a stacked autoencoder to extract damage-related features from the aligned signal, and cluster the extracted features using spectral clustering. Construct a pseudo-label dataset based on the clustering results. S4. Use the pseudo-label dataset to train a convolutional neural network for the classification process to classify the unlabeled single-sample damaged terahertz dataset. S5. Perform terahertz category-coded imaging on the classification results to obtain high-resolution and high-contrast two-dimensional images of the damage.
[0009] Furthermore, in step S1, a reflective terahertz time-domain spectroscopy system with a bandwidth of 10 GHz to 4.5 THz is used to obtain the time-domain terahertz signal. The specific system consists of a femtosecond laser, an integrated terahertz transmitter and receiver, an optical transmission system, and a data acquisition and control system.
[0010] Furthermore, in step S2, the data alignment layer is constructed based on the physical prior characteristics of the first pulse of the terahertz signal. By selecting a reference terahertz signal and using the unit impulse response function as the signal alignment function, the terahertz signal to be processed is aligned through convolution operation.
[0011] Furthermore, in step S3, the stacked autoencoder adopts a layered architecture, which includes two processes: encoding and decoding. It mainly consists of an encoding layer, a bottleneck layer, and a decoding layer, and is used to extract damage-related features from the aligned terahertz signal.
[0012] Furthermore, the encoding layer of the stacked autoencoder contains multiple sub-layers, which achieve progressive dimensionality compression through linear transformation and ReLU activation function. The bottleneck layer outputs key feature vectors that are highly correlated with the damage. The decoding layer achieves dimensionality reconstruction through a mirror encoding process and minimizes the root mean square error between the reconstructed signal and the original signal.
[0013] Furthermore, in step S3, the spectral clustering algorithm constructs and normalizes the Laplacian matrix by calculating the similarity matrix of the feature vectors, maps the data to a low-dimensional space to preserve the data structure characteristics, and performs K-Means clustering to output cluster labels to distinguish between damaged and normal regions.
[0014] Furthermore, in step S4, the convolutional neural network includes convolutional layers, max-pooling layers, ReLU non-linear activation layers, and fully connected layers. It is trained using a pseudo-labeled dataset, and the parameters are updated using the Adam optimizer. The optimal model parameters are obtained using the cross-entropy loss function.
[0015] in, This is the probability of predicting a damaged area.
[0016] Furthermore, in step S5, the terahertz category-coded imaging method involves constructing a label matrix and an imaging matrix, mapping the classification labels to specific pixel values, and achieving high-resolution and high-contrast imaging of damaged and normal areas.
[0017] Furthermore, the composite laminate samples are prepared using 3D printing technology. Each sample contains multiple layers of glass fiber reinforced polymer material, and air gaps of different thicknesses are introduced at the center of the bottom of the sample to simulate delamination defects.
[0018] Compared with the prior art, the beneficial effects of the present invention are: First, existing deep learning methods in terahertz damage detection heavily rely on large amounts of manually labeled data, resulting in high costs and long processing times. This invention, however, utilizes an unsupervised clustering strategy to generate pseudo-labels, enabling the construction of a training dataset for the classification network without manual annotation. This transforms unlabeled signals into effective training data, completely eliminating the dependence on labeled data and better meeting the practical needs of industrial sites where data is scarce. Second, traditional methods fail to consider the impact of terahertz signals on operating conditions such as lifting distance and surface uniformity, leading to a sharp drop in model performance and insufficient generalization when faced with new samples or changing test conditions. This invention addresses this by using a signal alignment layer based on terahertz physical priors, achieving signal synchronization with the peak value of the first pulse as a benchmark. This fundamentally reduces distribution shifts caused by differences in signals, significantly improving the model's adaptability to different samples and test conditions, and solving the problem of weak generalization ability in traditional models. Furthermore, traditional terahertz imaging relies on manually selected features, and damage features in complex multi-layered structures are easily submerged by noise, resulting in low imaging resolution and blurred boundaries. The proposed class-coded imaging method directly constructs images using the classification results of convolutional neural networks as pixel values, eliminating the need for manual feature selection. Through significant contrast of binary labels, damage boundaries are clearly discernible, resulting in superior imaging accuracy and efficiency compared to traditional methods. In particular, while existing technologies require large-scale, multi-source labeled datasets to guarantee model performance, this invention significantly reduces the dependence on data volume. This "trainable from a single sample" characteristic enables rapid response to the detection needs of industrial sites, facilitating practical engineering applications.
[0019] In summary, this invention, through a two-stage unsupervised learning framework design, physics-driven signal preprocessing, and high-resolution imaging strategy, has achieved improvements and breakthroughs over existing technologies in terms of label dependence, generalization ability, imaging accuracy, and engineering practicality, providing support for the industrial application of terahertz nondestructive testing technology in composite material damage identification. Attached Figure Description
[0020] Figure 1 This is the overall flowchart of the present invention; Figure 2 A schematic diagram illustrating the working principle of the experimental equipment of this invention; Figure 3 These are sample images prepared under different damage thicknesses according to the present invention; Figure 4 The diagram shows the terahertz signal to be processed and the terahertz signal after alignment processing in this invention. Figure 5 This is a structural diagram of the stack autoencoder of the present invention; Figure 6 This is a two-dimensional imaging map of the damage caused by clustering results in this invention; Figure 7 This is a diagram of the convolutional neural network structure of the present invention; Figure 8 This is a two-dimensional imaging image of the damage resulting from the classification of this invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. The present invention will be further described with reference to the accompanying drawings and embodiments: like Figures 1-8 As shown, the terahertz imaging method for composite material damage based on two-stage unsupervised learning includes the following steps: S1. Prepare composite laminate samples with different damage thicknesses, obtain time-domain terahertz signals of different samples based on terahertz time-domain spectroscopy system, and construct multiple unlabeled single-sample damage terahertz datasets. To characterize the damage within the composite material, the terahertz time-domain spectroscopy system employs a bandwidth of 10 GHz to 4.5 THz to obtain time-domain terahertz signals. The terahertz time-domain spectroscopy system includes a femtosecond laser, an integrated terahertz transmitter and receiver, an optical transmission system, and a data acquisition and control system. Its system principle is as follows: Figure 2 As shown.
[0022] The specific principle is as follows: a femtosecond laser generates an ultrashort pulse beam, which is split into two paths by a beam splitter: a pump pulse and a probe pulse. The pump pulse excites a terahertz emitter, generating a terahertz pulse, which is then guided to the sample via an optical transmission system. The reflected terahertz pulse is collected and focused onto a terahertz receiver. Simultaneously, the probe pulse, through an adjustable optical delay line, allows for precise control of the time delay between the terahertz pulse and the probe pulse. This enables the receiver to record the electric field intensity of the terahertz pulse over time and convert the optical signal into an electrical signal, which is then synchronously acquired by a data acquisition system. By scanning the delay line, the complete time-domain waveform of the terahertz pulse can be reconstructed, processed by a lock-in amplifier, and stored in a computer, providing corresponding terahertz data for subsequent analysis.
[0023] like Figure 3As shown, to acquire terahertz (THz) signals to validate the proposed method, glass fiber reinforced composite laminate samples were specifically designed and fabricated using 3D printing technology. These samples consisted of square composite laminates, each with a 10-layer structure, made by stacking glass fiber reinforced polymer (GFRP) layers, with each sample measuring 20 mm × 20 mm × 2 mm. To simulate delamination defects of different thicknesses, defects of varying thicknesses were inserted into the bottom of the four samples. Specifically, the defect thicknesses were 100 μm, 400 μm, 700 μm, and 1000 μm, corresponding to the four training samples. These defects were introduced during the 3D printing process in the form of controlled air gaps to ensure consistency in geometric and physical properties between samples, while controlling the air gap thickness to generate samples with different defect thicknesses.
[0024] Furthermore, based on the aforementioned THz-TDS system, to verify the effectiveness and generalization ability of the proposed framework on limited unlabeled terahertz data, a dedicated unlabeled single-sample dataset was constructed using the four prepared samples. Specifically, the THz signals of the four samples were acquired by point-by-point scanning using the THz-TDS system to construct four single-sample datasets. During the detection process, signals were acquired sequentially from the four samples at a scan step size of 0.5 mm. At this point, each tested sample contained 1600 sampling points, corresponding to 1600 THz response signals from the damaged and normal regions, respectively. The ratio of the number of THz signals acquired from the damaged region to the number of signals acquired from the normal region was approximately 1:5.
[0025] S2. Construct a data alignment layer based on the prior of the terahertz signal to perform data alignment processing on the acquired terahertz time-domain signal; During terahertz detection, the terahertz signal waveform changes primarily due to two key factors: first, the difference in lift distance (the perpendicular distance between the terahertz probe and the sample surface) between different samples; and second, the difference in surface uniformity at different sampling points within the same sample (inconsistent local thickness due to sample roughness or surface damage). During testing, these changes not only cause fluctuations in the terahertz signal but also directly affect subsequent feature extraction and model generalization performance. Specifically, when measuring different samples, even small changes in lift distance alter the propagation path of the terahertz wave, directly leading to time shifts, amplitude attenuation, and pulse width broadening in the first pulse, thus causing changes in the terahertz signal distribution across different samples. Even for the same sample, differences in surface uniformity due to surface roughness and defects in different test areas affect the reflection path of the terahertz wave, causing time and amplitude shifts in the first terahertz pulse. Both of these factors can lead to significant changes in the first terahertz signal pulse, such as… Figure 4As shown in (a) and (c), the differences between terahertz signals caused by lifting distance and surface uniformity are illustrated, respectively. It is noteworthy that the first pulse of the terahertz signal, being the part with the most concentrated signal energy, is directly affected by temporal position deviations, which alter the overall signal distribution characteristics. This difference impacts the deep learning model's ability to learn effective impairment features, particularly affecting the consistency of gradient calculations during backpropagation, leading to a decrease in model generalization ability. Therefore, constructing a data alignment layer to align terahertz signals in the temporal domain is a crucial step in improving model performance.
[0026] Furthermore, in step S2, the data alignment layer is constructed based on the physical prior characteristics of the first pulse of the terahertz signal. By selecting a reference terahertz signal and using the unit impulse response function as the signal alignment function, the terahertz signal to be processed is aligned through convolution operation.
[0027] This invention constructs a data alignment layer based on the physical prior characteristics of the first pulse of a terahertz signal. Specifically, a reference terahertz signal is selected, the position of the first pulse of the terahertz signal is fixed as a reference, and the unit impulse response function is used as the signal alignment function. The terahertz signal to be processed is aligned through a convolution operation. The specific alignment process can be defined as follows: (1) in, For the terahertz signal to be processed, For the aligned terahertz signal, This represents the time difference between the first pulse delay of the terahertz signal to be processed and the reference terahertz signal.
[0028] In particular, since the alignment operation process is consistent with the convolution operation process in deep learning, a data alignment layer can be built in the model design to perform data alignment processing.
[0029] Furthermore, for the input terahertz signal, through the processing of the data alignment layer, different terahertz signals can be aligned in the time domain, such as... Figure 4 As shown, the terahertz signals before and after alignment are displayed. Figure 4 (a) and (c) show the reference terahertz signal before alignment and the changed terahertz signal. It can be found that there is a significant deviation between the terahertz signals, which is due to the difference in the lift distance when measuring different samples. Figure 4 (a) or the difference in surface uniformity at different sampling points ( Figure 4 (c) causes a change in the propagation path of the terahertz signal, which directly leads to a significant shift in the timing of the first pulse. Figure 4(b) and (d) show the aligned terahertz signals. It can be observed that the terahertz signals exhibit significant consistency in the first pulse. This is achieved by applying an alignment layer to precisely synchronize the first pulse of the signal to be processed with the reference signal, significantly improving the overall waveform similarity. This alignment operation ensures that the first pulse, where the signal energy is most concentrated, achieves temporal structure consistency, enabling the model to achieve stable learning of damage features and laying a solid foundation for improving the model's generalization ability.
[0030] S3. Construct a two-stage unsupervised learning framework, including clustering and classification processes. Use a stacked autoencoder to extract damage-related features from the aligned signal, and cluster the extracted features using spectral clustering. Construct a pseudo-label dataset based on the clustering results. Following the alignment operation, a two-stage unsupervised learning framework was specifically constructed for damage identification. This framework includes unsupervised clustering and classification processes. The unsupervised clustering process extracts damage-related features from the aligned terahertz signals and generates a pseudo-label dataset, providing a data source for the subsequent classification process. The classification process classifies the terahertz signals, and the classification results can be used for subsequent high-resolution damage imaging.
[0031] Furthermore, the unsupervised clustering process is introduced. Its overall process mainly consists of three parts: feature extraction, cluster analysis, and pseudo-label dataset construction. The specific process is as follows: First, a stacked autoencoder is used to extract features from the terahertz signal, which is crucial for obtaining damage-related distinguishable features. Traditional unsupervised clustering methods typically extract distinguishable features directly from the original input signal, but their feature extraction capability for complex signals is limited, which may degrade model performance. To further improve clustering performance, a stacked autoencoder is used to extract damage-related features from the aligned terahertz signal before clustering.
[0032] like Figure 3 As shown, in step S3, the stacked autoencoder adopts a layered architecture, which includes two processes: encoding and decoding. It mainly consists of an encoding layer, a bottleneck layer, and a decoding layer, and is used to extract damage-related features from the aligned terahertz signal.
[0033] The encoding layer of the stacked autoencoder contains multiple sub-layers. It achieves progressive dimensionality compression through linear transformation and ReLU activation function. The bottleneck layer outputs key feature vectors that are highly correlated with the impairment. The decoding layer reconstructs the dimension through a mirror encoding process and minimizes the root mean square error between the reconstructed signal and the original signal.
[0034] Specifically, the coding layer contains four sub-layers, which achieve progressive dimensionality compression through linear transformation and the ReLU activation function. This is applicable to terahertz signals with a dimension of 2048. XAfter input, the data is processed through four encoding sub-layers, compressing the output feature dimension to 16. The output feature of each layer can be defined as: (2) in, For the first l -1 layer input, and The first l Layer weights and biases.
[0035] The bottleneck layer, as the core connection between the encoding and decoding layers, is used to directly output the 16-dimensional feature vector. F It retains key information highly correlated with damage in the terahertz signal, removes redundant noise, and provides effective input for subsequent clustering. The decoding layer also contains four sub-layers, achieving dimensional reconstruction through a mirror encoding process, that is, using linear transformation and the ReLU activation function to reconstruct the 16-dimensional features. F The decoding process is as follows: (3) in, and They represent the first l The weights and biases of the decoding sublayers, and This represents the output of the sub-layer, used to reconstruct intermediate features. The final result is a 2048-dimensional reconstructed signal. The decoding process is achieved by minimizing X and The root mean square error (RMSE) ensures the output of the bottleneck layer. F It can accurately reflect the damage characteristics of the input signal.
[0036] Through this gradual dimensionality reduction process, the hierarchical features of the input terahertz signal are extracted step by step, ultimately compressing the input into a 16-dimensional representation in the bottleneck layer. The bottleneck layer, containing 16 neurons and a bias term, is the core structure and highest coding layer for feature extraction, preserving the damage-related features in the terahertz signal, which is crucial for subsequent defect identification tasks. The decoding layer also contains four sub-layers, achieving dimensionality reconstruction through a mirror coding process. It uses linear transformation and the ReLU activation function to transform the 16-dimensional features... F By sequentially restoring the signal, a 2048-dimensional reconstructed signal is finally obtained. The decoding process is achieved by minimizing X and The root mean square error (RMSE) ensures the output of the bottleneck layer. F It can accurately reflect the damage characteristics of the original signal.
[0037] In step S3, the spectral clustering algorithm constructs and normalizes the Laplacian matrix by calculating the similarity matrix of the feature vectors, maps the data to a low-dimensional space to preserve the data structure characteristics, and performs K-Means clustering to output cluster labels to distinguish between damaged and normal regions.
[0038] Because the damaged region in a composite material structure is very limited compared to the normal region, the damaged terahertz signal set often exhibits sparsity and imbalance. This invention employs spectral clustering for clustering. As a graph-based clustering method, spectral clustering aims to minimize the connectivity between different clusters by partitioning the graph structure. In this case, each data point can be considered a node in the graph, and the similarity matrix between different data points can be used to calculate the corresponding edge weights. Subsequently, a normalized Laplacian matrix can be constructed to map the data to a low-dimensional space and capture the structural characteristics of the graph. The spectral clustering algorithm relies solely on the similarity matrix between data points, requiring no prior knowledge of the data distribution. Therefore, it is more suitable for handling clusters of arbitrary shapes, especially when dealing with imbalanced clusters, where it has a significant advantage over traditional clustering algorithms.
[0039] In this case, the input to the spectral clustering algorithm is the feature vector obtained by the encoder, and the output is the clustering result. The specific process is as follows: for a given input feature... F Calculate the similarity matrix; construct and normalize the Laplacian matrix based on the similarity matrix to map high-dimensional features to a low-dimensional space to preserve data structure characteristics; finally, perform K-Means clustering on the low-dimensional data and output the cluster label P (0 or 1, where 0 represents the predicted damaged area and 1 represents the predicted normal area), which can be defined as: (4) in, This represents the spectral clustering operation. In this study, the number of clusters is represented. (Corresponding to damaged and normal categories). For all input terahertz signals, the final output of spectral clustering can be expressed as: (5) Finally, based on the spectral clustering results, a pseudo-label dataset is constructed as the data required for subsequent classification. Based on the clustering labels, the clustering results of the samples are imaged using category encoding imaging in S5, such as... Figure 6 As shown in the diagram. Predicted normal areas are shown in yellow with a cluster label of 1, predicted damaged areas are shown in blue with a cluster label of 0, and actual damaged areas are marked with red boxes. Figure 6 As can be seen, the clustering results can initially distinguish between damaged and normal regions. Damaged regions are located in the center of the sample, while normal regions are located on the periphery. Signals with a cluster label of 0 are selected from the center of the blue region (where the damage confidence is high). and assign pseudo-labels This represents the damaged terahertz signal. Signals with a cluster label of 1 are selected from the outer edge of the yellow area (where the confidence level is high). and assign pseudo-labels This represents the undamaged terahertz signal in the normal region. The filtered signals are combined with the corresponding pseudo-labels to form a pseudo-labeled terahertz dataset, which is used to train the convolutional neural network in the subsequent classification process.
[0040] S4. Use the pseudo-label dataset to train a convolutional neural network for the classification process to classify the unlabeled single-sample damaged terahertz dataset. Following the unsupervised clustering process in step S3, a convolutional neural network (CNN) is trained on the pseudo-labeled dataset obtained in step S3 to obtain the optimal classification model. This trained CNN is then used to classify the unlabeled single-sample damaged terahertz dataset, outputting the classification label for each sampling point. The specific process is as follows: First, in the terahertz signal classification process, to reduce model complexity and simplify design, a classic convolutional neural network is adopted, the structure of which is as follows: Figure 7 As shown.
[0041] The convolutional neural network includes convolutional layers, max pooling layers, ReLU non-linear activation layers, and fully connected layers. It is trained using a pseudo-labeled dataset, and the parameters are updated using the Adam optimizer. The optimal model parameters are obtained using the cross-entropy loss function.
[0042] Specifically, the network comprises two convolutional layers, two max-pooling layers with a stride of 2, two ReLU nonlinear activation layers, and two fully connected layers. During network execution, the pseudo-labeled dataset is input into the network, and features are extracted through convolution and pooling, then integrated by the fully connected layers. The final output is the probability distribution of damaged and normal regions, along with the corresponding binary classification labels (0 or 1). Notably, the use of a simple CNN architecture for the classification process not only verifies that the proposed framework does not rely on complex model architecture design, simplifying the framework's complexity, but also demonstrates good damage recognition performance and computational efficiency. This is beneficial for promoting the application of this framework in practical terahertz nondestructive testing.
[0043] Secondly, a classification network was trained using a pseudo-label dataset. The pseudo-label dataset was divided into training and validation sets in an 8:2 ratio for training and testing. The Adam optimizer was used to update the parameters, and the cross-entropy loss function was employed to obtain the optimal model parameters. (6) in, This is the probability of predicting a damaged area.
[0044] Finally, after training is completed, the unlabeled single-sample damaged terahertz dataset is sequentially input into the trained model to classify each sample data and obtain the prediction result for each sampling point of the corresponding sample. Finally, it is represented by the classification label 0 or 1, which represent the damaged signal and the non-damaged signal, respectively.
[0045] S5. Perform terahertz category-coded imaging on the classification results to obtain high-resolution and high-contrast two-dimensional images of the damage; In terahertz imaging, traditional time-domain and frequency-domain methods usually rely on manual selection of imaging features (such as amplitude, peak position, time delay difference, etc.). However, for complex multilayer composite structures, damage features are easily hidden in the terahertz signal, and the best imaging features are difficult to identify accurately, resulting in low imaging resolution and contrast, or even damage artifacts.
[0046] To address this, the present invention proposes a novel category-coded terahertz high-resolution imaging method, the specific imaging process of which is as follows: First, a label matrix is constructed. Based on the obtained classification labels, the sampling points are arranged sequentially according to their spatial coordinates to form a two-dimensional label matrix that matches the physical size of the sample. The value of each element in the matrix is the classification label at that location (each sampling point corresponds to label 0 or 1, where 0 marks the damaged area and 1 marks the normal area).
[0047] Secondly, an imaging matrix is constructed. To achieve the mapping between labels and pixel values, a binary encoding rule is used to map label 0 (damaged area) to a specific pixel value (represented as blue in the image), and label 1 (normal area) to another pixel value (represented as yellow in the image), thereby constructing an imaging matrix to achieve high-resolution imaging of the damage. In particular, the difference in image color can intuitively distinguish between damaged and normal areas.
[0048] like Figure 6The image shows the two-dimensional imaging result after spectral clustering in this invention, i.e., the visualization image of the clustering result obtained based on step S5. From left to right, the image presents the clustering imaging effect of four samples with damage thicknesses of 100 μm, 400 μm, 700 μm, and 1000 μm. Damaged areas are marked in blue, normal areas are marked in yellow, and the actual damage range is marked by a red box. Analysis shows that the overall imaging accuracy after clustering is low, and it decreases as the defect thickness decreases: at a defect thickness of 1000 μm, the blue area completely covers the actual damaged area, with clear boundaries and no significant misclassification, resulting in the best clustering effect; when the defect thickness is 700 μm, a small number of yellow areas appear inside the damaged area (i.e., the damage is misclassified as normal), while there are sporadic blue dots in the normal area (i.e., misclassified as damage); for a defect with a thickness of 400 μm, a continuous blue cluster area is formed in the central region, but the continuous normal areas in the upper left and upper right corners of the sample are misclassified as damage; and a defect with a thickness of 100 μm shows a large area of misclassification in the composite material sample, indicating that the clustering method is difficult to effectively identify thinner defects.
[0049] like Figure 8 As shown, from left to right, the class-coded imaging results after the S4 step are presented for four samples with different damage thicknesses (100 μm, 400 μm, 700 μm, and 1000 μm). The distribution characteristics of the damage area are clearly presented through grayscale differences. The defect areas in the four sub-images are clearly and completely detected, with well-defined boundaries that are highly consistent with the actual defect range. The internal and external boundaries of the defects are clearly distinguishable, providing an accurate representation of the damage location, shape, and size. The leftmost image shows the 100 μm damage sample. The blue area (damage) in the image is located at the center of the sample. Although the damage thickness is the smallest, the boundary is complete and discernible, without the feature submersion phenomenon commonly seen in traditional imaging. This indicates that the classification network captures the terahertz signal changes caused by subtle damage, and the class-coded imaging effectively preserves the contour information of the damage. In addition, there is no significant noise interference in the defect area or background area of all samples, ensuring high-fidelity visualization of the defects. This demonstrates that our framework can effectively extract and utilize the discriminative features of terahertz signals, accurately locate defects, and generate clear and reliable defect images, outperforming traditional terahertz imaging methods in terms of defect detection integrity and boundary clarity.
[0050] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of protection claimed by the present invention. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A terahertz imaging method for composite material damage based on two-stage unsupervised learning, characterized in that, Includes the following steps: S1. Prepare composite laminate samples with different damage thicknesses, obtain time-domain terahertz signals of different samples based on terahertz time-domain spectroscopy system, and construct multiple unlabeled single-sample damage terahertz datasets. S2. Construct a data alignment layer based on the prior of the terahertz signal to perform data alignment processing on the acquired terahertz time-domain signal; S3. Construct a two-stage unsupervised learning framework, including clustering and classification processes. Use a stacked autoencoder to extract damage-related features from the aligned signal, and cluster the extracted features using spectral clustering. Construct a pseudo-label dataset based on the clustering results. S4. Use the pseudo-label dataset to train a convolutional neural network for the classification process to classify the unlabeled single-sample damaged terahertz dataset. S5. Perform terahertz category-coded imaging on the classification results to obtain high-resolution and high-contrast two-dimensional images of the damage.
2. The composite material damage terahertz imaging method based on two-stage unsupervised learning according to claim 1, characterized in that, In step S1, the terahertz time-domain spectroscopy system uses a bandwidth of 10 GHz to 4.5 THz to obtain time-domain terahertz signals. The terahertz time-domain spectroscopy system includes a femtosecond laser, an integrated terahertz transmitter and receiver, an optical transmission system, and a data acquisition and control system.
3. The composite material damage terahertz imaging method based on two-stage unsupervised learning according to claim 1, characterized in that, In step S2, the data alignment layer is constructed based on the physical prior characteristics of the first pulse of the terahertz signal. By selecting a reference terahertz signal and using the unit impulse response function as the signal alignment function, the terahertz signal to be processed is aligned through convolution operation.
4. The composite material damage terahertz imaging method based on two-stage unsupervised learning according to claim 1, characterized in that, In step S3, the stacked autoencoder adopts a layered architecture, which includes two processes: encoding and decoding. It mainly consists of an encoding layer, a bottleneck layer, and a decoding layer, and is used to extract damage-related features from the aligned terahertz signal.
5. The composite material damage terahertz imaging method based on two-stage unsupervised learning according to claim 4, characterized in that, The encoding layer of the stacked autoencoder contains multiple sub-layers. It achieves progressive dimensionality compression through linear transformation and ReLU activation function. The bottleneck layer outputs key feature vectors that are highly correlated with the impairment. The decoding layer reconstructs the dimension through a mirror encoding process and minimizes the root mean square error between the reconstructed signal and the original signal.
6. The composite material damage terahertz imaging method based on two-stage unsupervised learning according to claim 1, characterized in that, In step S3, the spectral clustering algorithm constructs and normalizes the Laplacian matrix by calculating the similarity matrix of the feature vectors, maps the data to a low-dimensional space to preserve the data structure characteristics, and performs K-Means clustering to output cluster labels to distinguish between damaged and normal regions.
7. The composite material damage terahertz imaging method based on two-stage unsupervised learning according to claim 1, characterized in that, In step S4, the convolutional neural network consists of convolutional layers, max-pooling layers, ReLU non-linear activation layers, and fully connected layers. It is trained using a pseudo-labeled dataset, and the parameters are updated using the Adam optimizer. The optimal model parameters are obtained using the cross-entropy loss function. ; in, This is the probability of predicting a damaged area.
8. The composite material damage terahertz imaging method based on two-stage unsupervised learning according to claim 1, characterized in that, In step S5, the terahertz category-coded imaging method maps classification labels to specific pixel values by constructing a label matrix and an imaging matrix, thereby achieving high-resolution and high-contrast imaging of damaged and normal areas.
9. The composite material damage terahertz imaging method based on two-stage unsupervised learning according to claim 1, characterized in that, The composite laminate samples were prepared using 3D printing technology. Each sample contained multiple layers of glass fiber reinforced polymer material, and air gaps of different thicknesses were introduced at the center of the bottom of the sample to simulate delamination defects.
Citation Information
Cited By
Damage detection method, electronic equipment, storage medium and computer program product
CN121431428A