Carbon steel eddy current thermal imaging corrosion detection method based on space-time attention self-coding

By adopting a spatiotemporal and spatial attention self-coding method in early corrosion detection, the STAAE model combined with LSTM and self-attention mechanisms is used to solve the problems of DCAE in dynamic temperature feature learning, noise impact and information loss, and achieve higher detection sensitivity and accuracy.

CN120232802APending Publication Date: 2025-07-01NANJING TECH UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510379715.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing deep convolutional autoencoder (DCAE) has problems in early corrosion detection with insufficient ability to learn dynamic temperature characteristics, susceptibility to noise, and information loss during feature encoding, resulting in insufficient detection sensitivity and accuracy.

Method used

Using a spatiotemporal attention autoencoding method, the spatiotemporal attention automatic encoder (STAAE), combined with a long and short-term memory network (LSTM) and a self-attention mechanism, capture the spatiotemporal dynamic changes of temperature data, and identify and visualize the corrosion area through the maximum reconstruction error calculation method.

Benefits of technology

It improves sensitivity to early corrosion signals, enhances visualization of corrosion areas, optimizes detection stability and accuracy, and can capture dynamic heat flow characteristics more accurately and improves detection contrast.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120232802A_ABST
    Figure CN120232802A_ABST
Patent Text Reader

Abstract

The invention discloses a carbon steel eddy current thermal imaging corrosion detection method based on space-time attention self-coding, and relates to the technical field of eddy current thermal imaging detection.The method comprises the steps that an early-stage corrosion sample is detected through pulsed eddy current thermal imaging, an infrared image sequence is obtained, and the change trend of the surface temperature of the early-stage corrosion sample along with time is captured; constructing a space-time attention automatic encoder, and training the space-time attention automatic encoder by using the infrared image sequence of the non-corrosion sample; taking the to-be-detected infrared image as a new input signal, capturing the spatial-temporal dynamic change of the temperature data by using the spatial-temporal attention automatic encoder, and generating a reconstruction signal; and constructing an error graph by calculating the maximum reconstruction error of the input signal and the reconstruction signal, and identifying and visualizing the corrosion area in the infrared image to be detected. The method is suitable for the conditions of weak early corrosion signals and fuzzy boundaries, and provides more reliable technical support for accurate identification of early corrosion areas and subsequent corrosion degree evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of eddy current thermography detection, and particularly to a carbon steel eddy current thermography corrosion detection method based on spatio-temporal attention auto-encoder. Background Art

[0002] Pulsed eddy current thermography (ECPT) technology actively excites and heats the surface of a conductive material, and collects an infrared image sequence to obtain the corrosion information of the material. However, in early corrosion detection, due to the small change in the heat conduction characteristics of the corrosion area, the local temperature difference is not obvious, and the contrast between the corrosion area and the non-corrosion area in the infrared image is low, making it difficult to effectively identify the early corrosion area. Traditional corrosion detection methods usually cannot make full use of the spatio-temporal characteristics of the temperature field, resulting in the difficulty of accurately characterizing the temperature evolution pattern of the corrosion area, thereby reducing the detection sensitivity.

[0003] On the other hand, existing deep learning methods have certain limitations in capturing the temperature dynamic changes of the corrosion area. The deep convolutional auto-encoder (DCAE) mainly relies on local spatial feature extraction and is difficult to effectively model the time series features of the corrosion area, resulting in the difficulty of accurately reconstructing the temperature evolution pattern of the corrosion area. In addition, traditional error calculation methods usually use global statistical indicators such as mean square error (MSE). These methods are prone to averaging errors when calculating the abnormality degree of the corrosion area, masking weak corrosion signals and reducing the detection accuracy of the early corrosion area.

[0004] In addition, during the ECPT detection process, the infrared thermal image sequence is easily interfered by factors such as environmental noise, heating non-uniformity, and background thermal reflection, further reducing the accuracy of feature extraction of the corrosion area. Existing auto-encoder methods fail to distinguish the importance of features at different time steps during the time series modeling process, resulting in insufficient attention of the model to the corrosion area when reconstructing the temperature distribution and affecting the visualization effect of the corrosion area.

[0005] In summary, the existing deep convolutional auto-encoder (DCAE) method has obvious limitations in early corrosion detection, which are mainly reflected in the following three aspects:

[0006] First, the Deep Convolutional Autoencoder (DCAE) has difficulty effectively capturing the dynamic temperature evolution characteristics of the corrosion area, resulting in limited accuracy in extracting corrosion signals. Theoretically, the Deep Convolutional Autoencoder (DCAE) mainly relies on convolutional layers to extract and reconstruct spatial features. However, its encoding mechanism mainly focuses on local spatial patterns and is difficult to effectively learn the time-varying heat diffusion trend. In practical applications, early corrosion signals often manifest as weak temperature gradient changes. Their temperature evolution process is affected by factors such as heat conduction, local oxidation, and material decay, showing strong time dependence. However, due to the lack of dedicated time modeling capabilities in the Deep Convolutional Autoencoder (DCAE), it is difficult to capture the dynamic response characteristics of the corrosion area during the heating-cooling cycle, resulting in a lower sensitivity of the reconstructed thermal image to changes in the temperature distribution of the corrosion area. In addition, the heat diffusion characteristics of the corrosion area may exhibit complex local dynamic behaviors due to the inhomogeneity of the material microstructure. The limitations of the Deep Convolutional Autoencoder (DCAE) in static feature extraction further reduce its applicability in corrosion detection tasks.

[0007] Second, the reconstruction mechanism of the Deep Convolutional Autoencoder (DCAE) is easily interfered by background noise, affecting the contrast and boundary clarity of the corrosion area. In infrared thermography detection, environmental temperature fluctuations, background thermal radiation, and detection device noise will all affect the stability of the temperature field. Especially in the early stage of corrosion, the temperature change in the corrosion area is small, and its signal is easily masked by noise. When the Deep Convolutional Autoencoder (DCAE) performs feature extraction and reconstruction, although it can learn the overall distribution of the image, it mainly relies on loss functions such as Mean Squared Error (MSE) for training, resulting in its tendency to regard background noise as important features and learn them together during the reconstruction process, thus affecting the reconstruction quality. Specifically, the contrast of the corrosion area in the reconstructed image decreases, the boundary becomes blurred, and even false features appear, thus affecting the accurate identification of the corrosion area. In addition, due to the failure of the Deep Convolutional Autoencoder (DCAE) to introduce an explicit noise suppression mechanism, the stability of its reconstruction results to environmental factors is poor, making the visualization effect of the corrosion area dependent on the quality of data preprocessing, further limiting the reliability of its practical application.

[0008] Finally, the deep convolutional autoencoder (DCAE) is prone to losing key information during the high-dimensional data compression and feature extraction processes, reducing the stability of detection. Infrared thermal imaging data usually contains rich spatio-temporal information. The temperature evolution pattern in the corrosion area not only has spatial distribution characteristics but also involves complex time series characteristics. However, during the feature encoding process of the deep convolutional autoencoder (DCAE), multiple convolutional layers and pooling layers are usually used for step-by-step dimensionality reduction. Although this compression mechanism can extract the main features, it is also easy to lose key information, especially the subtle temperature differences between the corrosion area and the background. In practical applications, the corrosion area usually only accounts for a small part of the image. Compared with the background area, its signal is relatively weak. The deep convolutional autoencoder (DCAE) often tends to focus on the global structure during feature learning and ignores these key local features, resulting in the final reconstructed corrosion area being not fine enough and even possibly causing the loss of some corrosion information. In addition, the feature expression ability of the deep convolutional autoencoder (DCAE) is greatly affected by network structure hyperparameters (such as convolutional kernel size, pooling method), resulting in limited generalization ability in different task scenarios and poor adaptability of the model under different corrosion degrees, materials, and detection conditions, further affecting its practicality.

[0009] Based on the above analysis, the limitations of the existing deep convolutional autoencoder (DCAE) in early corrosion detection are mainly reflected in the insufficient ability to learn dynamic temperature features, the susceptibility to noise leading to a decrease in visualization quality, and information loss during the feature encoding process. These technical defects make the deep convolutional autoencoder (DCAE) show instability in extracting corrosion signals, high sensitivity to noise, and limitations in the visualization effect of the corrosion area in practical applications, and it is difficult to meet the requirements of early corrosion detection for weak signal enhancement and accurate feature extraction. Summary of the Invention

[0010] Based on this, it is necessary to provide a carbon steel eddy current thermographic corrosion detection method based on spatio-temporal attention autoencoding for the above technical problems.

[0011] The present invention provides a carbon steel eddy current thermographic corrosion detection method based on spatio-temporal attention autoencoding, including:

[0012] S1. Using pulsed eddy current thermography to detect early corrosion samples, obtaining an infrared image sequence, and capturing the change trend of the surface temperature of the early corrosion samples over time;

[0013] S2. Constructing a spatio-temporal attention autoencoder and training the spatio-temporal attention autoencoder using the infrared image sequence of non-corroded samples;

[0014] S3. Taking the infrared image to be detected as a new input signal, using the spatio-temporal attention autoencoder to capture the spatio-temporal dynamic changes of the temperature data and generating a reconstructed signal;

[0015] S4. By calculating the maximum reconstruction error between the input signal and the reconstructed signal, construct an error map to identify and visualize the corrosion areas in the infrared image to be detected.

[0016] Furthermore, use pulsed eddy current thermography to detect early corrosion samples, obtain a sequence of infrared images, and capture the change trend of the surface temperature of the early corrosion samples over time, including:

[0017] S11. Obtain infrared images of several consecutive frames through pulsed eddy current thermography. Each frame of the infrared image represents the temperature distribution of the early corrosion sample at a time step;

[0018] S12. Represent the pixel values in the infrared image as the temperature values at the corresponding positions on the surface of the early corrosion sample, and smooth the pixel values of each frame of the infrared image to remove irregular noise and interference;

[0019] S13. Arrange the smoothed infrared images in chronological order to form an infrared image sequence in the form of time series data, and obtain the change trend of the surface temperature of the early corrosion sample in the time series dimension.

[0020] Furthermore, construct a spatio-temporal attention autoencoder, and use the infrared image sequence of the non-corroded sample to train the spatio-temporal attention autoencoder, including:

[0021] S21. Obtain the infrared image sequence of the non-corroded sample as the training data;

[0022] S22. Use the training data as the input signal and the time structure of the reconstructed sequence as the output signal to create a spatio-temporal attention autoencoder architecture that includes an encoding component, a self-attention mechanism, and a decoding component;

[0023] S23. Set the hyperparameters of the spatio-temporal attention autoencoder, train the spatio-temporal attention autoencoder using the mean squared error loss, and optimize it using the adaptive moment estimation algorithm to obtain the trained and optimized spatio-temporal attention autoencoder.

[0024] Furthermore, create a spatio-temporal attention autoencoder architecture that includes an encoding component, a self-attention mechanism, and a decoding component, with the training data as the input and the time structure of the reconstructed sequence as the output, including:

[0025] S221. Create an encoding component that includes two layers of long short-term memory networks, a convolutional layer, and a pooling layer. First, pass the input signal to the two layers of long short-term memory networks to capture the time development of the temperature data, then use the convolutional layer to collect spatial features, and then minimize the feature dimension through the pooling layer;

[0026] S222. After the convolutional layer and pooling layer of the encoding component, introduce a self-attention mechanism to assign adaptive weights to each time step in the input signal to enhance temporal features;

[0027] S223. Create a decoding component that includes two consecutive deconvolution blocks and two long short-term memory networks, recreate high-level thermal features according to the original input of the input signal, and through time alignment, output reconstructed data that conforms to the input format to obtain the reconstructed signal of the reconstructed output.

[0028] Further, the long short-term memory network in the encoding component updates its respective hidden state and cell state using the state calculation formula at each time step to maintain the memory of past information, where the state calculation formula is:

[0029]

[0030] In the formula, l represents the long short-term memory network layer index; h t represents the hidden state at time step t; c t represents the cell state at time step t; x t represents the input signal at time step t.

[0031] Further, the calculation formula for introducing the self-attention mechanism to assign adaptive weights to each time step in the input signal is:

[0032]

[0033] Q = W Q F2;

[0034] K = W K F2;

[0035] V = W V F2;

[0036] In the formula, Attention represents the self-attention mechanism; Q represents the query vector; K represents the key vector; V represents the value vector; softmax represents the normalized exponential function; F2 represents the encoded feature map; k represents the pooling scale; T represents the transpose matrix; d k represents the dimension of the query vector and the key vector; W Q , W K and W V all represent the learned weight matrices.

[0037] Further, the expression for the convolutional layer in the encoding component to collect spatial features is:

[0038] F l = ReLU(Conv1D(F l-1 ,Wl ) + b l );

[0039] In the formula, F l represents the result feature map after ReLU processing; F l-1 represents the input feature map of the previous layer; W l represents the convolution weight; b l represents the convolution bias.

[0040] Furthermore, the expression of the pooling layer in the encoding component is:

[0041] P l = MaxPooling1D(F l , k);

[0042] In the formula, P l represents the output feature map after pooling; k represents the pooling scale; F l represents the result feature map after ReLU processing; MaxPooling1D represents the one-dimensional maximum pooling operation.

[0043] Furthermore, by calculating the maximum reconstruction error between the input signal and the reconstructed signal, constructing an error map, and identifying and visualizing the corrosion area in the infrared image to be detected includes:

[0044] S41. Reshape the input signal and the reconstructed signal from the flat vector form into the original spatio-temporal dimensions, and respectively obtain the pixel intensities at any spatial position and any time step of the original frame and the reconstructed frame;

[0045] S42. Construct a pixel reconstruction error map at any time step by calculating the difference in pixel intensities between the original frame and the reconstructed frame;

[0046] S43. Select the maximum error value at each pixel position, summarize the pixel faults in the entire time dimension, construct a summarized error map, and visually display the corrosion area in the infrared image to be detected by showing the maximum reconstruction mismatch detected at each position point.

[0047] Furthermore, the construction formula of the pixel reconstruction error map is:

[0048] E t (i, j) = |X t (i, j) - X t (i, j)|;

[0049] In the formula, E t (i, j) represents the pixel reconstruction error map at time step t; X t (i, j) represents the pixel intensity of the original frame at spatial position (i, j) and time step t; X t(i, j) represents the pixel intensity of the reconstructed frame at the spatial position (i, j) and time step t;

[0050] The construction formula for the aggregated error map is:

[0051]

[0052] In the formula, E agg (i, j) represents the aggregated error map; T represents the number of frames.

[0053] The beneficial effects of the present invention are as follows:

[0054] 1. The present invention combines STAAE with ECPT, uses the STAAE model to learn the normal temperature evolution pattern of non-corroded samples, and accurately locates the corroded area through the reconstruction error during the detection process. Compared with the traditional Deep Convolutional Autoencoder (DCAE), STAAE can model the temporal dependence of temperature data in the time dimension, while DCAE mainly relies on convolutional layers to extract local spatial features and ignores the temporal variation law of corrosion signals. Therefore, the present invention uses an LSTM layer to capture time features and combines an adaptive attention mechanism to focus on key time steps, enabling the model to more accurately learn the dynamic thermal response of the corroded area and improving the sensitivity to early corrosion signals.

[0055] 2. The maximum reconstruction error calculation method proposed by the present invention can effectively avoid the problem that the mean error masks weak corrosion signals. The traditional DCAE method usually uses the Mean Squared Error (MSE) as a metric in the reconstruction error analysis. This method is prone to averaging out small corrosion signals by the global error, resulting in the difficulty of accurately identifying early corroded areas. The present invention constructs a high-contrast error distribution map by calculating the maximum error between the input signal and the reconstructed signal, making the abnormal features of the corroded area more obvious, thereby optimizing the visualization effect of the corroded area. This method is particularly suitable for the situation where early corrosion signals are weak and the boundaries are blurred, making the corroded area show clearer temperature anomaly features in the infrared thermal image, improving the stability and accuracy of detection. It can not only accurately capture the dynamic heat flow characteristics of the corroded area but also effectively improve the contrast and stability of corrosion detection, providing more reliable technical support for the accurate identification of early corroded areas and the subsequent assessment of corrosion degree. Description of the Drawings

[0056] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0057] Figure 1It is a flowchart of the carbon steel eddy current thermographic corrosion detection method based on spatio-temporal attention auto-encoding according to an embodiment of the present invention;

[0058] Figure 2 It is a structural block diagram of a pulsed eddy current thermographic system according to an embodiment of the present invention;

[0059] Figure 3 It is the reconstructed error images using a deep convolutional auto-encoder according to an embodiment of the present invention (a) corrosion for 14 days (b) corrosion for 21 days;

[0060] Figure 4 It is a data processing flowchart according to an embodiment of the present invention;

[0061] Figure 5 It is an architecture diagram of a spatio-temporal attention auto-encoder (STAAE) according to an embodiment of the present invention;

[0062] Figure 6 It is the maximum reconstructed error images according to an embodiment of the present invention (a) corrosion for 14 days (b) corrosion for 21 days. Detailed implementation manners

[0063] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0064] Please refer to Figure 1 , and there is provided a carbon steel eddy current thermographic corrosion detection method based on spatio-temporal attention auto-encoding, including:

[0065] S1. Use pulsed eddy current thermography to detect early corrosion samples, obtain an infrared image sequence, and capture the change trend of the surface temperature of the early corrosion samples over time.

[0066] In the description of the present invention, using pulsed eddy current thermography to detect early corrosion samples, obtain an infrared image sequence, and capture the change trend of the surface temperature of the early corrosion samples over time includes:

[0067] S11. Obtain infrared images of several consecutive frames through pulsed eddy current thermography, and each infrared image represents the temperature distribution of the early corrosion samples at a time step.

[0068] Specifically, the input data used in the present invention are thermal signal data obtained through pulsed eddy current thermography (ECPT). The original form of these thermal signal data is 122 infrared images, and the size of each image is 640×480 pixels. During the ECPT detection process, the temperature change is continuous in time, and each image represents the temperature distribution of the sample at a time step (i.e., a time point). The pixel value in each image represents the temperature value at the corresponding position on the sample surface.

[0069] S12. Represent the pixel values in the infrared images as the temperature values at the corresponding positions on the surface of the early corrosion sample, and smooth the pixel values of each infrared image to remove irregular noise and interference.

[0070] Specifically, these infrared images are smoothed to remove irregular noise and interference. The specific processing method is to smooth the pixel values of each image, making the thermal signals in the image smoother and more stable, thereby reducing the influence of external interference on the detection of temperature changes. This step is particularly important for accurately capturing temperature changes, especially when detecting early corrosion. Since the temperature change in the corrosion area is weak, the denoising process can effectively improve the stability of the signal.

[0071] S13. Arrange the smoothed infrared images in chronological order to form an infrared image sequence in the form of time series data, and obtain the change trend of the surface temperature of the early corrosion sample in the time series dimension.

[0072] Specifically, the processed 122 images are arranged in chronological order to form a time series data. After each image is smoothed, it constitutes a continuous thermal signal sequence, representing the temperature response of the sample over time. The size of these time series data is 307,200×1, where 307,200 is the total number of pixels of 122 images (640×480), and 1 represents single-channel data (i.e., the temperature value of each pixel).

[0073] In the processing of input data, the temperature change of each pixel point in the infrared image shows dynamic characteristics over time. By processing these image sequences, the change trend of the surface temperature of the sample can be captured in the time series dimension, and then the formation process of the corrosion area can be inferred. These input data provide the original data source for the subsequent feature extraction and analysis of the model, and are the basis for realizing efficient corrosion detection and positioning.

[0074] S2. Construct a spatio-temporal attention autoencoder, and use the infrared image sequence of non-corroded samples to train the spatio-temporal attention autoencoder.

[0075] In the description of the present invention, constructing a spatio-temporal attention autoencoder and using the infrared image sequence of non-corroded samples to train the spatio-temporal attention autoencoder includes:

[0076] S21. Obtain the infrared image sequence of the non-corroded sample as the training data.

[0077] S22. Using the training data as the input signal and the time structure of the reconstructed sequence as the output signal, create a spatio-temporal attention autoencoder architecture that includes an encoding component, a self-attention mechanism, and a decoding component.

[0078] Specifically, the present invention uses the thermal signal sequence of the non-corroded sample as the input to train the proposed spatio-temporal attention autoencoder (STAAE). The structure of the spatio-temporal attention autoencoder (STAAE) is as Figure 5 shown.

[0079] In the description of the present invention, creating a spatio-temporal attention autoencoder architecture that includes an encoding component, a self-attention mechanism, and a decoding component with the training data as the input and the time structure of the reconstructed sequence as the output includes:

[0080] S221. Create an encoding component that includes two layers of long short-term memory networks, a convolutional layer, and a pooling layer. First, pass the input signal to the two layers of long short-term memory networks to capture the temporal development of the temperature data, then use the convolutional layer to collect spatial features, and then minimize the feature dimension through the pooling layer.

[0081] Specifically, as Figure 5 shown, the encoding component includes two long short-term memory networks ( Figure 5 two LSTM Cell modules in Figure 5 ), a convolutional layer ( Figure 5 Conv1 in

[0082] ), and a pooling layer ( t c t and updating its hidden state h

[0083] Next, the spatio-temporal attention autoencoder uses a one-dimensional convolutional layer to collect spatial features and then performs a pooling operation. Each convolutional layer applies a one-dimensional kernel to the input sequence to create a feature map. After the convolutional stage, a one-dimensional max pooling layer is used to minimize the feature dimension and emphasize the most important spatial features.

[0084] In the description of the present invention, the long short-term memory network in the encoding component updates its respective hidden state and cell state using a state calculation formula at each time step to maintain the memory of past information. Among them, the state calculation formula is:

[0085]

[0086] In the formula, l represents the long short-term memory network layer index; h t represents the hidden state at time step t; c t represents the cell state at time step t; x t represents the input signal at time step t.

[0087] In the description of the present invention, the expression for the convolutional layer in the encoding component to collect spatial features is:

[0088] F l = ReLU(Conv1D(F l-1 , W l ) + b l );

[0089] In the formula, F l represents the resulting feature map after ReLU processing; F l-1 represents the input feature map of the previous layer; W l represents the convolutional weight; b l represents the convolutional bias.

[0090] In the description of the present invention, the expression for the pooling layer in the encoding component is:

[0091] P l = MaxPooling1D(F l , k);

[0092] In the formula, P l represents the output feature map after pooling; k represents the pooling scale; F l represents the resulting feature map after ReLU processing; MaxPooling1D represents a one-dimensional maximum pooling operation.

[0093] S222. After the convolutional layer and the pooling layer of the encoding component, a self-attention mechanism is introduced to assign adaptive weights to each time step in the input signal to enhance the temporal features.

[0094] Specifically, after convolution and pooling, a self-attention mechanism ([[]] Figure 5Attention in the model). It enhances temporal features by assigning adaptive weights to each time step. The self-attention mechanism calculates the attention weights by measuring the similarity between the feature representations of the query (Q), key (K), and value vector (V). The query (Q), key (K), and value vector (V) all come from the encoded feature map F2. The introduction of the self-attention mechanism ensures that the model focuses on the most relevant time part of the input sequence, thereby improving the ability to detect small temperature changes.

[0095] In the description of the present invention, a self-attention mechanism is introduced, and the calculation formula for assigning adaptive weights to each time step in the input signal is:

[0096]

[0097] Q=W Q F2;

[0098] K=W K F2;

[0099] V=W V F2;

[0100] In the formula, Attention represents the self-attention mechanism; Q represents the query vector; K represents the key vector; V represents the value vector; softmax represents the normalized exponential function; F2 represents the encoded feature map; k represents the pooling scale; T represents the transposed matrix; d k W represents the dimension of query vector and key vector; Q , W K and W V Both represent the learned weight matrices.

[0101] S223. Create a decoding component including two consecutive deconvolution blocks and two long short-term memory networks, recreate the high-level thermal features according to the original input of the input signal, and after time alignment, output the reconstructed data that conforms to the input format to obtain a reconstructed signal of the reconstructed output.

[0102] Specifically, the decoding component uses two consecutive deconvolution blocks ( Figure 5 Dconv1 and Dconv2) and two long short-term memory (LSTM) networks ( Figure 5 The two LSTM cells in the lower right corner of the middle) recreate high-level hot features as closely as possible according to the original input. Each deconvolution block imitates the Conv1D procedure, but uses deconvolution to upsample and restore the original feature dimension.

[0103] These LSTM layers then decipher the temporal structure of the reconstructed sequence.

[0104] Finally, the output part ensures that the reconstructed data conforms to the input format. This alignment is done by the time-distributed dense layer. It applies a shared weight matrix W out and bias b out . Therefore, the output format is consistent with the input format. The final reconstructed output X is obtained as follows:

[0105] X = TimeDistributed(Dense(F decoded ));

[0106] where F decoded represents the output result of the decoding component; TimeDistributed represents the time-distributed wrapper layer for processing sequence data.

[0107] S23. Set the hyperparameters of the spatio-temporal attention autoencoder, train the spatio-temporal attention autoencoder using the mean squared error loss, and optimize it using the adaptive moment estimation algorithm to obtain the trained and optimized spatio-temporal attention autoencoder.

[0108] Specifically, the spatio-temporal attention autoencoder is trained using the mean squared error (MSE) loss and optimized using the Adam algorithm (adaptive moment estimation algorithm). Before starting the training, the hyperparameters of the spatio-temporal attention autoencoder, such as the number of iterations and the learning rate, need to be determined. Training the autoencoder (AE) entirely on non-corrupted signals shows that it can reliably recover the normal mode. This technique can minimize the reconstruction error of the corrosion signal and contribute to early corrosion diagnosis. This technique completes feature extraction by calculating the difference between the reconstructed signal and the input signal and can be used to detect early corrosion areas. Subsequently, the network parameters are initialized and then optimized repeatedly. The number of epochs, batch size, and learning rate are set to 50, 100, and 0.001, respectively. The formula for the mean squared error (MSE) is:

[0109]

[0110] where N represents the number of samples; X i represents the true value of the i-th sample; X i represents the reconstructed output of the i-th sample.

[0111] After training the STAAE entirely on non-corroded (normal) thermal signal sequences, the model can reproduce these normal modes excellently with low error. When a new sequence that may involve corrosion is introduced, the model tries to reconstruct the sequence based on the learned normal data representation.

[0112] S3. Use the infrared image to be detected as a new input signal, and use the spatio-temporal attention autoencoder to capture the spatio-temporal dynamic changes of the temperature data and generate a reconstructed signal.

[0113] S4. Construct an error map by calculating the maximum reconstruction error between the input signal and the reconstructed signal, and identify and visualize the corrosion areas in the infrared image to be detected.

[0114] In the description of the present invention, constructing an error map by calculating the maximum reconstruction error between the input signal and the reconstructed signal, and identifying and visualizing the corrosion areas in the infrared image to be detected includes:

[0115] S41. Reshape the input signal and the reconstructed signal from the flat vector form into the original spatio-temporal dimension, and obtain the pixel intensities of the original frame and the reconstructed frame at any spatial position and any time step.

[0116] Specifically, the deviation between the input sequence and the reconstructed output will show the abnormal positions. To show the corrosion areas, the present invention first reshapes the reconstructed output X and the original input X from the flat vector form into the original spatio-temporal dimension. Let X t (i,j) and X t (i,j) represent the pixel intensities of the original frame and the reconstructed frame at the spatial position (i, j) and the time step t respectively.

[0117] S42. Construct a pixel reconstruction error map at any time step by calculating the difference in pixel intensities between the original frame and the reconstructed frame.

[0118] In the description of the present invention, the construction formula of the pixel reconstruction error map is:

[0119] E t (i,j) = |X t (i,j) - X t (i,j)|;

[0120] In the formula, E t (i, j) represents the pixel reconstruction error map at the time step t; X t (i,j) represents the pixel intensity of the original frame at the spatial position (i, j) and the time step t; X t (i,j) represents the pixel intensity of the reconstructed frame at the spatial position (i, j) and the time step t.

[0121] This pixel error highlights the positions where the model fails to accurately reproduce the expected normal thermal pattern. Since the model only learns normal conditions during training, this difference may be related to early corrosion or other abnormalities. To further improve the visibility of the defect metrics, the present invention aggregates the pixel faults over the entire time dimension by selecting the maximum error value at each pixel position. This technique highlights the most severe deviations in the entire series, effectively emphasizing the most severe corrosion symptoms. Formally, given a series of T frames, a summary error map E agg .

[0122] S43. Select the maximum error value at each pixel position, summarize the pixel faults over the entire time dimension, construct a summarized error map, and visualize the corrosion area in the infrared image to be detected by showing the maximum reconstruction mismatch detected at each position point.

[0123] In the description of the present invention, the construction formula of the summarized error map is:

[0124]

[0125] In the formula, E agg (i, j) represents the summarized error map; T represents the number of frames.

[0126] This maximum error aggregation method ensures that subtle but important changes are not masked by averaging. Instead, the generated error map clearly shows the maximum reconstruction mismatch detected at each position point, providing a clear and concise image of the possible corrosion area. In E agg The high-intensity areas in the figure correspond to the places where the reconstruction of the normal mode of the model fails most severely, thus effectively identifying the early corrosion areas with higher sensitivity and accuracy.

[0127] By combining an LSTM layer for time-dependent modeling, a convolutional pooling layer for spatial feature extraction, and a self-attention mechanism for adaptive time focusing, the proposed model can efficiently capture the spatio-temporal dynamic features of temperature data. By testing different filters, number of units, and adjusting the layer depth, the optimal design of the model is shown in Table 1.

[0128] Table 1: Specific parameters of the spatio-temporal attention autoencoder (STAAE)

[0129]

[0130] Based on the above analysis, the model (spatio-temporal attention autoencoder) proposed by the present invention is trained using non-corroded samples, and the trained model is used to detect the corrosion area. By calculating the maximum reconstruction error with the original input image, the maximum reconstruction error image is finally obtained. The present invention can effectively reconstruct the features of the corrosion area, thereby accurately locating the early corrosion area.

[0131] Using the proposed spatio-temporal attention autoencoder to detect the corrosion samples on the 14th day ( Figure 6 a) in and the 21st day ( Figure 6 b) in, the maximum reconstruction error image is obtained. It can be seen from Figure 6 that the thermal images generated by this model are not only clearer and more obvious, but also can more accurately express the potential corrosion pattern, so it surpasses the traditional methods in terms of clarity and diagnostic reliability.

[0132] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0133] The present invention adopts pulsed eddy current thermography (ECPT) technology, and the system structure is as Figure 2 shown.

[0134] The working process includes three stages:

[0135] (1) Transient thermal excitation stage: A transient high-frequency excitation current is applied to the surface of the carbon steel to be measured through an electromagnetic excitation module. Based on the skin effect, an eddy current field is generated on the surface layer of the material, and a non-uniform Joule heat distribution is caused due to the difference in conductor resistivity;

[0136] (2) Thermal response acquisition stage: During the natural cooling process after the thermal excitation stops, the dynamic attenuation process of the temperature field on the surface of the material is synchronously acquired through a high-frame-rate infrared thermal imager, generating thermal image time-series data with spatio-temporal continuity;

[0137] (3) Feature analysis stage: The computer processes the acquired image sequence, uses different algorithms to obtain corrosion-related feature parameters, and accordingly performs corrosion image reconstruction and corrosion degree evaluation.

[0138] Currently, the main deep learning algorithm for corrosion detection using pulsed eddy current thermography technology is the deep convolutional autoencoder. The steps of using the deep convolutional autoencoder can be divided into:

[0139] ① Feature extraction

[0140] Regarding the infrared image sequence as a temperature sampling sequence that changes over time, it is input into the deep convolutional autoencoder (DCAE) for automatic feature learning, extracting the high-level spatio-temporal features of each pixel point in the image to form a low-dimensional latent representation.

[0141] Z = f enc (X; θ) = σ(W l *X + b l );

[0142] Among them, X represents the input infrared image sequence, Z is the low-dimensional latent feature after being mapped by the encoder, f enc is the encoder, which uses convolution or pooling layers to map the input infrared image sequence X to the low-dimensional latent feature representation Z; W l , b l are the convolution kernel weights and biases of the encoder; θ is the training parameter of the neural network; σ is the non-linear activation function; the encoder uses convolution operations to extract spatial features and combines time series information at the same time to enhance the temperature difference between the defective area and the non-defective area.

[0143] ② Reconstruct the feature image

[0144] Use the decoder to reconstruct the low-dimensional features to obtain a denoised thermal image, enhance the contrast of the defect area, and improve the visualization effect:

[0145] X = f dec (Z; φ) = σ(W l ′ * Z + b l ′);

[0146] where f dec is the decoder, which uses transposed convolution (deconvolution) or upsampling to restore the original image size; W l ′, b l ′ are the parameters of the decoder; X is the reconstructed image.

[0147] ③ Generate the defect visualization result

[0148] Based on the reconstruction error or specific feature channels, generate the final defect thermal image to highlight the feature information of the defect area. The reconstruction error image is used to identify abnormal areas, reflect the thermal response intensity of the defect, and enhance the visibility of the defect.

[0149] The reconstruction error images of the above-mentioned deep convolutional autoencoder are as Figure 3 shown. a is the reconstruction error map after 14 days of corrosion, and b is the reconstruction error map after 21 days of corrosion.

[0150] The spatio-temporal attention autoencoder (STAAE) proposed in the present invention aims to detect and visualize the early atmospheric corrosion of steel structures. The model is trained using the thermal signal sequences of uncorroded samples. This enables it to learn normal spatio-temporal patterns. When applied to corroded samples, the model reconstructs their thermal signals. Then, the maximum reconstruction error of these signals is used to create an error image to help identify and locate the corrosion areas. STAAE integrates the self-attention mechanism with the LSTM layer and the deep convolutional autoencoder architecture.

[0151] The overall logical route is as Figure 4 shown, which shows the overall process of the proposed corrosion detection method. First, an electromagnetic field is generated by coil excitation and acts on the sample to be tested. During the eddy current heating process, an infrared thermal imager collects the thermal image sequence of the sample surface in real time. The thermal image sequence collected from the non-corroded area is used as training data and input into the spatio-temporal attention autoencoder model. The LSTM layer is used to model the time dependence, the convolutional-pooling module extracts the spatial features, and the self-attention mechanism is used to strengthen the key temporal information to complete the learning and modeling of the normal thermal response.

[0152] In the detection stage, a sequence of thermal images from the corroded area is input into the pre-trained autoencoder. Since the model only learns the thermal features of the non-corroded area, it cannot accurately reconstruct the abnormal thermal response, resulting in significant reconstruction errors in the corroded area. By visualizing the error map between the input image and the reconstructed image, the maximum reconstruction error map is further extracted as the corrosion feature map to achieve the localization and detection of the early corroded area.

[0153] Aiming at the technical problems existing in the application of the existing deep convolutional autoencoder (DCAE) in corrosion detection, such as insufficient feature extraction, inadequate utilization of reconstruction errors, and low sensitivity to weak corrosion signals, the present invention proposes a corrosion detection method combining electromagnetic induction pulsed thermography (ECPT) technology with spatio-temporal attention autoencoder (STAAE), and innovatively introduces the calculation of the maximum reconstruction error to optimize the visualization of the corroded area, having the following advantages:

[0154] First, the present invention combines STAAE with ECPT, and uses the STAAE model to learn the normal temperature evolution pattern of non-corroded samples, and accurately locates the corroded area through the reconstruction error during the detection process. Compared with the traditional deep convolutional autoencoder (DCAE), STAAE can model the temporal dependence of temperature data in the time dimension, while DCAE mainly relies on convolutional layers to extract local spatial features, ignoring the temporal variation law of corrosion signals. Therefore, the present invention uses an LSTM layer to capture temporal features and combines an adaptive attention mechanism to focus on key time steps, enabling the model to more accurately learn the dynamic thermal response of the corroded area and improving the sensitivity to early corrosion signals.

[0155] Second, the maximum reconstruction error calculation method proposed by the present invention can effectively avoid the problem that the mean error masks weak corrosion signals. The traditional DCAE method usually uses the mean square error (MSE) as a metric in the reconstruction error analysis. This method easily averages out small corrosion signals by the global error, resulting in the difficulty of accurately identifying early corroded areas. The present invention constructs a high-contrast error distribution map by calculating the maximum error between the input signal and the reconstructed signal, making the abnormal features of the corroded area more obvious, thereby optimizing the visualization effect of the corroded area. This method is particularly suitable for the situation where early corrosion signals are weak and the boundaries are blurred, making the corroded area show clearer temperature anomaly features in the infrared thermal image, improving the stability and accuracy of detection.

[0156] By combining STAAE with ECPT and introducing the calculation of the maximum reconstruction error, the present invention breaks through the limitations of the existing DCAE method in spatio-temporal feature extraction, visualization of corrosion areas, and sensitivity to early signals. This method can not only accurately capture the dynamic heat flow characteristics of the corrosion area, but also effectively improve the contrast and stability of corrosion detection, providing more reliable technical support for the accurate identification of early corrosion areas and the subsequent assessment of corrosion degree.

[0157] In addition, time series modeling methods based on other principles can replace some algorithm steps in the present invention, but the present invention has obvious advantages in terms of method innovation and pertinence. At present, the methods of time series modeling are not limited to long short-term memory networks (LSTM), and methods such as Transformer architecture, gated recurrent unit (GRU), or variational autoencoder (VAE) can also be used; similarly, the application methods of feature learning and attention mechanism of time series can also be replaced by time series Transformer (Time-Series Transformer) or self-supervised learning-based methods. However, the present invention adopts the method of combining STAAE with ECPT, uses the LSTM layer to model the temporal dependence of temperature signals, and focuses on key time steps through an adaptive attention mechanism, ensuring that the model can more accurately learn the dynamic heat response of the corrosion area and improving the sensitivity to early corrosion signals.

[0158] Compared with the Transformer architecture, LSTM has stronger stability in processing small sample time series data and can avoid the problem of high resource consumption during long time series dependence calculation. In addition, the adaptive attention mechanism of the present invention is optimized for the characteristics of ECPT thermal signals and can more effectively highlight the abnormal heat flow characteristics of the corrosion area, rather than simply relying on the global self-attention mechanism to weight the entire time series. Therefore, in the ECPT corrosion detection task, the method of the present invention has stronger pertinence and robustness, and can balance the accuracy and efficiency of temporal feature extraction on the premise of controllable computational complexity.

[0159] In summary, the innovation of the present invention lies in the deep integration of STAAE and ECPT, as well as the joint application of LSTM and the adaptive attention mechanism. Although methods such as Transformer can replace LSTM to a certain extent for time series modeling, no other technical approaches that can completely replace the solution of the present invention and achieve the same detection effect have been found yet.

[0160] In summary, by means of the above technical solutions of the present invention, STAAE is combined with ECPT. The STAAE model is used to learn the normal temperature evolution pattern of non-corroded samples, and the corroded area is accurately located through the reconstruction error during the detection process. Compared with the traditional deep convolutional autoencoder (DCAE), STAAE can model the temporal dependence of temperature data in the time dimension, while DCAE mainly relies on the convolutional layer to extract local spatial features and ignores the temporal variation law of corrosion signals. Therefore, the present invention uses an LSTM layer to capture time features and combines an adaptive attention mechanism to focus on key time steps, enabling the model to more accurately learn the dynamic thermal response of the corroded area and improving the sensitivity to early corrosion signals. The maximum reconstruction error calculation method proposed in the present invention can effectively avoid the problem that the mean error masks weak corrosion signals. When analyzing the reconstruction error, the traditional DCAE method usually uses the mean square error (MSE) as a metric, which is likely to average out small corrosion signals by the global error, making it difficult to accurately identify early corroded areas. By calculating the maximum error between the input signal and the reconstructed signal, the present invention constructs a high-contrast error distribution map, making the abnormal features of the corroded area more obvious, thereby optimizing the visualization effect of the corroded area. This method is particularly suitable for cases where early corrosion signals are weak and the boundaries are blurred, making the corroded area show more distinct temperature anomaly features in the infrared thermal image, improving the stability and accuracy of detection. It can not only accurately capture the dynamic heat flow characteristics of the corroded area, but also effectively improve the contrast and stability of corrosion detection, providing more reliable technical support for the accurate identification of early corroded areas and the subsequent evaluation of corrosion degree.

[0161] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially according to the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

Claims

1. A carbon steel eddy current thermal imaging corrosion detection method based on spatiotemporal attention self-encoding, characterized in that: include: S1. Use pulsed eddy current thermal imaging to detect early corrosion samples, obtain infrared image sequences, and capture the change trend of surface temperature of early corrosion samples over time; S2. Construct a spatiotemporal attention autoencoder and train it using infrared image sequences without corrosion samples. S3, taking the infrared image to be detected as a new input signal, using the spatiotemporal attention autoencoder to capture the spatiotemporal dynamic changes of the temperature data and generate a reconstructed signal; S4. By calculating the maximum reconstruction error between the input signal and the reconstructed signal, an error map is constructed to identify and visualize the corrosion area in the infrared image to be detected.

2. The carbon steel eddy current thermal imaging corrosion detection method based on spatiotemporal attention self-encoding according to claim 1 is characterized in that: The method of using pulsed eddy current thermal imaging to detect early corrosion samples, obtain infrared image sequences, and capture the trend of surface temperature changes of early corrosion samples over time includes: S11, obtaining a plurality of continuous frames of infrared images by pulsed eddy current thermal imaging, each frame of infrared image representing the temperature distribution of the early corrosion sample in one time step; S12, using the pixel values ​​in the infrared image to represent the temperature value of the corresponding position on the surface of the early corrosion sample, and smoothing the pixel values ​​of each frame of the infrared image to remove irregular noise and interference; S13, arranging the smoothed infrared images in chronological order to form an infrared image sequence in the form of time series data, and obtaining a change trend of the surface temperature of the early corrosion samples in the time series dimension.

3. The carbon steel eddy current thermal imaging corrosion detection method based on spatiotemporal attention self-encoding according to claim 1 is characterized in that: The construction of the spatiotemporal attention autoencoder and the training of the spatiotemporal attention autoencoder using the infrared image sequence of the non-corroded sample include: S21, obtaining an infrared image sequence of a non-corroded sample as training data; S22. Using the training data as input and the temporal structure of the reconstructed sequence as output, a spatiotemporal attention autoencoder architecture is created, which includes an encoding component, a self-attention mechanism, and a decoding component. S23. Set the hyperparameters of the spatiotemporal attention autoencoder, train the spatiotemporal attention autoencoder using the mean square error loss, and optimize it using the adaptive moment estimation algorithm to obtain the trained and optimized spatiotemporal attention autoencoder.

4. The carbon steel eddy current thermal imaging corrosion detection method based on spatiotemporal attention self-encoding according to claim 3 is characterized in that: The architecture of the spatiotemporal attention autoencoder, which takes the training data as input and reconstructs the temporal structure of the sequence as output, and creates an encoding component, a self-attention mechanism, and a decoding component, includes: S221, creating an encoding component including a two-layer long short-term memory network, a convolution layer, and a pooling layer, first passing the input signal to the two-layer long short-term memory network to capture the temporal development of the temperature data, then using the convolution layer to collect spatial features, and then minimizing the feature dimension through the pooling layer; S222, after the convolutional layer and pooling layer of the encoding component, a self-attention mechanism is introduced to assign adaptive weights to each time step in the input signal to enhance the temporal features; S223. Create a decoding component including two consecutive deconvolution blocks and two long short-term memory networks, recreate the high-level thermal features according to the original input of the input signal, and after time alignment, output the reconstructed data that conforms to the input format to obtain a reconstructed signal of the reconstructed output.

5. The carbon steel eddy current thermal imaging corrosion detection method based on spatiotemporal attention self-encoding according to claim 4 is characterized in that: The long short-term memory network in the encoding component uses a state calculation formula to update the respective hidden state and cell state at each time step to maintain the memory of past information, wherein the state calculation formula is: Where l represents the long short-term memory network layer index; h t represents the hidden state at time step t; c t represents the cell state at time step t; x t represents the input signal at time step t.

6. The carbon steel eddy current thermal imaging corrosion detection method based on spatiotemporal attention self-encoding according to claim 4 is characterized in that: The expression of collecting spatial features by the convolutional layer in the encoding component is: F l =ReLU(Conv1D(F l-1 ,IN l )+b l ); In the formula, F l Represents the result feature map after ReLU processing; F l-1 Represents the input feature map of the previous layer; W l represents the convolution weight; b l Represents the convolution bias.

7. The carbon steel eddy current thermal imaging corrosion detection method based on spatiotemporal attention self-encoding according to claim 4 is characterized in that: The expression of the pooling layer in the encoding component is: P l =MaxPooling1D(F l ,k); Where P l represents the output feature map after pooling; k represents the pooling scale; F l Represents the resulting feature map after ReLU processing; MaxPooling1D represents a one-dimensional maximum pooling operation.

8. The carbon steel eddy current thermal imaging corrosion detection method based on spatiotemporal attention self-encoding according to claim 4 is characterized in that: The self-attention mechanism is introduced to calculate the adaptive weight assigned to each time step in the input signal: Q=W Q F2; K=W K F2; V=W V F2; In the formula, Attention represents the self-attention mechanism; Q represents the query vector; K represents the key vector; V represents the value vector; softmax represents the normalized exponential function; F2 represents the encoded feature map; k represents the pooling scale; T represents the transposed matrix; d k Represents the dimension of query vector and key vector; W Q , W K and W V Both represent the learned weight matrices.

9. The carbon steel eddy current thermal imaging corrosion detection method based on spatiotemporal attention self-encoding according to claim 1 is characterized in that: The method of calculating the maximum reconstruction error between the input signal and the reconstructed signal, constructing an error map, and identifying and visualizing the corrosion area in the infrared image to be detected includes: S41, reshape the input signal and the reconstructed signal from the flat vector form into the original space-time dimension, and obtain the pixel intensity of the original frame and the reconstructed frame at any spatial position and any time step respectively; S42, constructing a pixel reconstruction error map at any time step by calculating the pixel intensity difference between the original frame and the reconstructed frame; S43. Select the maximum error value at each pixel position, summarize the pixel failures in the entire time dimension, construct a summary error map, and visualize the corrosion area in the infrared image to be detected by displaying the maximum reconstruction mismatch detected at each position point.

10. The carbon steel eddy current thermal imaging corrosion detection method based on spatiotemporal attention self-encoding according to claim 9 is characterized in that: The construction formula of the pixel reconstruction error map is: E t (i,j)=|X t (i,j)-X t (i,j)|; In the formula, E t (i, j) represents the pixel reconstruction error map at time step t; X t (i, j) represents the pixel intensity of the original frame at spatial position (i, j) and time step t; X t (i, j) represents the pixel intensity of the reconstructed frame at spatial position (i, j) and time step t; The construction formula of the summary error map is: In the formula, E agg (i, j) represents the summary error map; T represents the frame number.

Citation Information

Cited By

  • Compressed thermal infrared image temperature calculation method based on improved Transform architecture

    CN121095732A