Unsupervised small sample internal defect detection method
By combining a pulsed thermal imaging system with a deep autoencoder and the Swin Transformer Wnet model for unsupervised learning, the problems of high cost, low efficiency and insufficient accuracy in internal defect detection in intelligent manufacturing are solved, and high-precision unsupervised internal defect detection is achieved.
Patent Information
- Application Number
- CN202511000799.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-14
AI Technical Summary
Existing methods for detecting internal defects in the field of intelligent manufacturing suffer from high costs, low efficiency, and limited applicability. In particular, unsupervised learning methods are prone to incomplete noise suppression, blurred defect boundaries, and limited detection accuracy when processing complex thermal imaging data, making it difficult to meet the requirements of industrial applications.
An unsupervised method for detecting internal defects in small samples is constructed by acquiring infrared thermal images using a pulsed thermal imaging system and conducting unsupervised learning through a deep autoencoder and a Swin Transformer Wnet model. This method includes using a deep autoencoder to compress temporal thermal images to generate latent spatial features, using a Swin Transformer Wnet for feature extraction and reconstruction, and combining fully connected conditional random fields for post-processing to achieve defect segmentation.
High-precision internal defect detection was achieved even with scarce and unlabeled defect samples, solving the problems of small sample size, high noise, and blurred boundaries, and providing a high-precision and low-cost internal defect detection solution.
Smart Images

Figure CN120953185A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent manufacturing technology, and in particular to an unsupervised method for detecting internal defects in small samples. Background Technology
[0002] In the field of intelligent manufacturing, the detection technology for internal defects in polymer materials is facing severe technical challenges. Although these materials are widely used in high-end manufacturing fields such as aerospace and new energy due to their excellent physicochemical properties, the presence of internal defects seriously threatens the mechanical properties and long-term reliability of the products.
[0003] The main technical bottleneck in the inspection process lies in the identification of internal defects. Current materials inspection technologies exhibit a clear polarization: surface inspection technologies have achieved industrial-grade applications, while internal inspection technologies are hampered by various factors. From a technical implementation perspective, existing internal defect detection methods generally have application limitations. For example, X-ray inspection faces cost issues, ultrasonic inspection is constrained by material properties, and magnetic particle and eddy current inspections have limited applicability. These methods fail to meet the requirements of advanced manufacturing in terms of inspection efficiency, economy, and safety. Infrared thermal imaging technology, due to its unique advantages, is considered a potential solution, but it also faces dual challenges in industrial applications: on the one hand, the nonlinear characteristics of thermal imaging data cause traditional analysis methods to fail; on the other hand, supervised learning-based intelligent inspection methods are limited by the difficulty of obtaining labeled data in industrial scenarios. While current unsupervised learning methods alleviate the data dependency problem to some extent, they still have significant shortcomings when processing complex thermal imaging data: incomplete noise suppression, blurred defect boundaries, and limited detection accuracy restrict their industrial application value.
[0004] For example, invention patent CN118797469A discloses a low-data-dependency defect detection method for industrial assembly scenarios based on the Swin Transformer model. This method includes: acquiring a large defect detection dataset containing both normal and defective samples from an industrial assembly scenario, and dividing the dataset into several subsets based on location information; pre-training a Swin Transformer network model using the large defect detection dataset to obtain a defect detection model; and fine-tuning the defect detection model using partial labeled data from a new scenario. However, this method relies on partial labeled data to fine-tune the defect detection model and identify defects.
[0005] For example, CN114841977A discloses a defect detection method based on a SwinTransformer structure combined with SSIM and GMSD. The method includes: acquiring an industrial defect image; inputting it into a pre-constructed SwinTransformer-based feature extraction network for feature learning; reconstructing the defect-free region information from the learned feature map to obtain a reconstructed image; calculating the structural similarity index (SSIM) between the industrial defect image and the reconstructed image to obtain a defect anomaly feature map based on structural similarity; calculating the gradient magnitude similarity deviation (GMSD) between the industrial defect image and the reconstructed image to obtain a defect anomaly feature map based on gradient magnitude similarity; and fusing the defect anomaly feature map based on structural similarity with the defect anomaly feature map based on gradient magnitude similarity to obtain the final defect detection result image. This method improves the accuracy and localization precision of defect detection, but it mainly focuses on surface defects and is difficult to apply to internal defect detection. Summary of the Invention
[0006] The purpose of this invention is to provide an unsupervised small-sample internal defect detection method. This method can effectively detect complex internal defect information of various products when defect samples are scarce and unlabeled, thereby promptly identifying production problems.
[0007] An embodiment provides a personalized load assessment feedback method that integrates user and task characteristics, comprising the following steps:
[0008] Step 1: Based on the pulsed thermal imaging system, the original infrared thermal image is acquired by adjusting the thermal imaging parameters of the sample to be tested, and after preprocessing, a time-series thermal image sequence of multiple defect samples is obtained.
[0009] Step 2: Input the time-series thermal image sequence into the deep autoencoder, and construct the total loss function of the deep autoencoder to train the deep autoencoder to obtain the trained deep autoencoder. The deep autoencoder includes: compressing the input time-series thermal image sequence through the encoder to generate latent spatial features, and reconstructing the latent spatial features into a denoised time-series thermal image sequence through the decoder.
[0010] Step 3: Input the denoised time-series thermal image sequence into the Swin Transformer Wnet, and construct the total loss function to train the Swin Transformer Wnet to obtain the trained Swin Transformer Wnet. The Swin Transformer Wnet includes: extracting features from the denoised time-series thermal image sequence through the encoder Swin Unet to generate a latent spatial image, and reconstructing the generated latent spatial image into the original input time-series thermal image sequence through the decoder Swin Unet.
[0011] Step 4: Input the new temporal thermal image sequence into the trained SwinTransformerWnet to generate the corresponding latent space image. Use a fully connected conditional random field for post-processing to obtain the final internal defect segmentation result.
[0012] In one embodiment, step 1, which involves acquiring the original infrared thermal image by adjusting the thermal imaging parameters of the sample to be tested, includes: adjusting the heating time, defect depth, and placement position of the sample to be tested in the pulsed thermal imaging system to acquire the original infrared thermal image.
[0013] Further, in step 1, the acquisition of time-series thermal image sequences of multiple defect samples after preprocessing includes: cropping the acquired original infrared thermal images, retaining the regions containing defect samples, and scaling the cropped regions using the resize method to obtain time-series thermal image sequences of multiple defect samples.
[0014] In one embodiment, in step 2, multiple hidden layers are stacked on top of the deep autoencoder. The encoder function and decoder function in the autoencoder of a single hidden layer are F(·) and G(·), respectively, and both the encoder function and decoder function in the autoencoder use the hyperbolic tangent function tanh as the activation function. The calculation formulas for the encoder function and decoder function are as follows:
[0015] z = F(x) = f(W1x + b1),
[0016] y = G(z) = "(W2z + b2),
[0017] In the formula, x represents the input temporal thermal image sequence, z represents the latent space representation generated by the encoder, y represents the reconstruction approximation generated by the decoder, W1 and b1 represent the weight matrix and bias of the encoder, respectively, and W2 and b2 represent the weight matrix and bias of the decoder, respectively.
[0018] In one embodiment, the loss function constructed in step 2 includes: using mean squared error to quantize the difference between the input and reconstructed output of the deep autoencoder, constraining the parameters of the deep autoencoder, and introducing an L2 regularization term to denoise the reconstructed approximation result. The calculation formula is as follows:
[0019] L total =L mse +L reg ,
[0020]
[0021] In the formula, L total Let L represent the total loss function. mseL represents the loss function constructed using the mean squared error. reg Let x represent the L2 regularization loss function, x represent the input time-series thermal image sequence, and y represent the reconstruction approximation result generated by the decoder. n Let y represent the time-series thermal image sequence input at the nth time point. n This represents the approximate reconstruction result at the nth time point generated by the decoder, where N represents the temporal length, λ represents the regularization coefficient, and w m This represents the m-th weight parameter in the decoder function, where M represents the total number of weight parameters.
[0022] In one embodiment, in step 3, the encoder Swin Unet and the decoder Swin Unet are cascaded to form a symmetrical structure. Both the encoder Swin Unet and the decoder Swin Unet include an image block partitioning submodule, a linear embedding submodule, a SwinTransformer submodule, an image block merging submodule, and a skip connection submodule.
[0023] The image block partitioning submodule is used to divide the denoised time-series thermal image sequence into multiple non-overlapping blocks, and flatten each block into a one-dimensional time-series thermal image sequence that matches the input requirements of the Swin Transformer submodule.
[0024] The linear embedding submodule is used to map the flattened cube into high-dimensional features;
[0025] The Swin Transformer submodule is used to generate a nonlinear feature image by taking the feature image output by two-dimensional convolution of multiple non-overlapping squares and high-dimensional features as input.
[0026] The image block merging submodule is used to downsample the nonlinear feature image and then gradually restore the resolution and detailed information of the feature image through upsampling.
[0027] The skip connection submodule is used to combine shallow features of the encoder Swin Unet with deep features of the decoder Swin Unet, perform linear transformations per feature channel using convolution, and obtain the latent spatial image of the encoder output using an activation function.
[0028] In one embodiment, the Swin Transformer submodule includes two consecutive subunits. Each subunit includes a layer normalization layer, a window-based multi-head self-attention layer, a normalization layer, and a multilayer perceptron layer. The window-based multi-head self-attention layer includes a window-based multi-head self-attention layer and a sliding window-based multi-head self-attention layer. The function expressions for the different layers are as follows:
[0029]
[0030] In the formula, x l-1 This represents the input characteristics of the Swing Transformer submodule, corresponding to the output of layer l-1; This represents the output features of the multi-head self-attention layer of the window. x represents the output feature of the multi-head self-attention layer of the sliding window. l and x l+1 Let LN represent the l-th and l+1-th output features of the multilayer perceptron layer, Norm represent the normalization operation, WMSA and SWMSA represent the multi-head self-attention operation of the window and the sliding window, respectively, and MLP represent the multilayer perceptron operation.
[0031] In one embodiment, downsampling based on a nonlinear feature image includes:
[0032] The nonlinear feature image is divided into multiple non-overlapping local windows. Pixels sharing the same spatial position on the local windows are connected along the channel dimension to form a feature tensor with half the length and width and four times the channel dimension. Then, application layer normalization is applied, and linear projection is performed to compress the channel dimension, generating a downsampling result with half the length and width and twice the channel dimension.
[0033] In one embodiment, the resolution and detailed information of the feature image are then gradually restored through upsampling, including:
[0034] Based on the downsampling results, a linear mapping is applied to double the number of feature channels, and the spatial resolution of the features is doubled. At the same time, the number of channels is reduced to one-quarter, gradually restoring the resolution and detailed information of the feature image.
[0035] In one embodiment, the total loss function of the Swin Transformer Wnet is constructed by a linear combination of the cross-entropy loss function, the normalized cut loss in graph theory, and the boundary open loss based on morphological operations.
[0036] The formula for calculating the cross-entropy loss function is as follows:
[0037]
[0038] In the formula, L ce Let y represent the cross-entropy loss function. c Let p represent the c-th real image. c This represents the predicted probability of the Swing TransformerWnet, where C represents the number of predicted classes.
[0039] The normalized cut loss in graph theory penalizes cross-subgraph connections by minimizing inter-class similarity and maximizing intra-class similarity, and the calculation formula is as follows:
[0040]
[0041] In the formula, Ncut D (G) represents the normalized cut loss function, A d Let G represent the set of pixels belonging to category d, where D represents the number of categories, and G represents the set of pixels belonging to all categories in the image. `cut(A` d GA d The assoc(A) is used to calculate the similarity weights between pixels in category d and pixels in other categories. d G) is used to calculate the similarity weight between pixels in category d and pixels in all categories. (u, v) represents the pixel similarity weight, which measures the similarity between pixel u and pixel v.
[0042] The similarity weights of pixels u and v are calculated using the following formula:
[0043]
[0044] In the formula, I(u) represents the gray intensity of pixel u, I(v) represents the gray intensity of pixel v, ||I(u)-I(v)|| represents the gray intensity difference, and σ I The grayscale similarity scale parameter is represented by X(u), where X(v) represents the spatial coordinates of pixel u, X(v) represents the spatial coordinates of pixel v, and ||X(u)-X(v)|| represents the spatial Euclidean distance. σ X The parameter represents the spatial similarity scale, and r represents the neighborhood radius, which is used to limit the calculation distance of the similarity weight;
[0045] The formula for calculating the boundary openness loss based on morphological operations is as follows:
[0046]
[0047] In the formula, L open Let H denote the boundary open loss function, H represent the number of pixels in the feature image, P represent the latent space image generated by the encoder Swin Unet, p represent the pixel location, and I(p) represent the original gray value of pixel p. This indicates a morphological etching operation. This represents the morphological dilation operation, and S represents the structuring element.
[0048] In one embodiment, step 4 uses fully connected conditional random field post-processing to obtain the final internal defect segmentation result, including: for the corresponding latent space image generated by Swin TransformerWnet, applying a fully connected conditional random field to achieve fine segmentation by minimizing the global energy function, the calculation formula is as follows:
[0049] E(α)=∑ i ψ u (α i )+∑ i<j ψ p (α i α j ),
[0050] In the formula, E(α) represents the global energy function of pixel label α, and ψ u Denotes a univariate potential function that causes the SwinTransformer Wnet to select labels with higher confidence, α i Let α represent the predicted label of the i-th pixel. j Let ψ represent the predicted label of the j-th pixel. p Let represent a binary potential function that penalizes similar pixels i and j by assigning them different labels.
[0051] Compared with the prior art, the beneficial effects of the present invention include at least the following:
[0052] (1) The present invention first constructs a deep autoencoder, and compresses the temporal thermal image sequence of the defect sample into a low-dimensional latent space through unsupervised defect detection, thereby improving the reconstruction accuracy of the original input image. It also constructs a corresponding loss function to preserve defect features, suppress thermal imaging noise, adapt to the temporal correlation of thermal image sequences, and effectively avoid the limitations of traditional supervised defect detection.
[0053] (2) Further based on the Swin Transformer, a Swin TransformerWnet model with a continuous encoding and decoding structure is constructed to generate a latent space image. Detection is achieved under zero-labeled samples. A corresponding loss function is constructed to optimize the image and the boundary accuracy is optimized. Pixel-level optimization is achieved by combining fully connected conditional random field post-processing. Finally, industrial-grade segmentation results are output, thereby achieving unsupervised accurate detection of internal defects under scarce sample conditions.
[0054] (3) This invention solves the three major technical bottlenecks of small sample size, high noise, and blurred boundaries in defect sample internal detection through unsupervised architecture design, multimodal feature optimization and industrial scenario adaptability, and provides a high-precision and low-cost sample internal defect detection solution for intelligent manufacturing. Attached Figure Description
[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0056] Figure 1 This is a flowchart illustrating the unsupervised small-sample internal defect detection method based on Swin Transformer provided by the present invention.
[0057] Figure 2 A comparison of the original thermal images provided for the embodiments and defect detection using the present invention. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and given in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0059] The solution of the present invention is as follows: Figure 1 As shown, an unsupervised small-sample internal defect detection method based on Swin Transformer includes the following steps:
[0060] Step 1: Based on the pulsed thermal imaging system, the original infrared thermal image is acquired by adjusting the thermal imaging parameters of the sample to be tested, and after preprocessing, a time-series thermal image sequence of multiple defect samples is obtained.
[0061] The pulsed thermal imaging setup includes a halogen lamp, a time-delay relay, an infrared camera, a lifting platform, and a computer. The time-delay relay generates a pulse signal to control the halogen lamp's heating of the defective sample. After heating, the infrared camera captures the surface temperature distribution of the defective sample during the cooling phase, and the computer stores the temperature data and displays the thermal images in real time. The lifting platform adjusts the relative position between the halogen lamp and the defective sample to ensure as uniform an illumination as possible.
[0062] Data Acquisition: A manually fabricated circular acrylic plate with internal defects was used. By adjusting the heating time, defect depth, and placement, richer thermal images of the temperature were obtained. Specifically, to obtain the surface temperature response under different total radiation energy conditions, eight different heating times were applied to the defective sample: 1 second, 1.5 seconds, 2 seconds, 2.5 seconds, 3 seconds, 3.5 seconds, 4 seconds, and 4.5 seconds. Defect depths were combinations of 1 mm and 2 mm, and combinations of 1.5 mm and 2.5 mm. Placement positions included both horizontal and vertical orientations. After heating, an infrared camera captured the cooling process of the sample at a frequency of 50 Hz for ten seconds, obtaining 500 images per experiment.
[0063] Data preprocessing: The original image resolution is 640*512. The images are cropped, leaving only square regions containing defective samples. The cropped regions of interest are scaled to a resolution of 200*200 using the `resize` method in Python. The dataset is randomly split in a 7:1 ratio, with 14,000 images used to train the proposed SwinTransformer Wnet, and the remaining 2,000 images used to test the trained SwinTransformer Wnet model.
[0064] S2. Input the time-series thermal image sequence into the deep autoencoder, and construct the total loss function of the deep autoencoder to train the deep autoencoder to obtain the trained deep autoencoder. The deep autoencoder includes: compressing the input time-series thermal image sequence through the encoder to generate latent spatial features, and reconstructing the latent spatial features into a denoised time-series thermal image sequence through the decoder.
[0065] In this embodiment, a one-dimensional deep autoencoder is trained using the original thermal image data to compress and reconstruct the time-series thermal image sequence, reducing nonlinear noise and retaining key defect features. The decoder output image of the deep autoencoder is used as the noise reduction result to establish the corresponding thermal image dataset.
[0066] Specifically, a deep autoencoder was developed to enhance the defect segmentation capability of SwinTransformerWnet for denoising preprocessed temporal thermal image sequences. This unsupervised neural network compresses the input data into a low-dimensional latent space through the encoder and reconstructs it through the decoder. For a basic autoencoder containing only a single hidden layer, its encoder function F(·) and decoder function G(·) are shown in equations (1) and (2):
[0067] z=F(x)=f(W1x+b1) (1)
[0068] y=G(z)=g(W2z+b2) (2)
[0069] In the formula, x represents the input temporal thermal image sequence, z represents the latent space representation generated by the encoder, y represents the reconstruction approximation generated by the decoder, W1 and b1 represent the weight matrix and bias of the encoder, respectively, and W2 and b2 represent the weight matrix and bias of the decoder, respectively.
[0070] The activation functions used by the encoder function F(·) and the decoder function G(·), namely the hyperbolic tangent function tanh, are used to introduce nonlinear transformation to process the nonlinear characteristics of thermal images. Their mathematical expression is shown in equation (3):
[0071]
[0072] The encoder module is used to compress the input temperature sequence data to remove redundant noise information, generate a corresponding compact feature representation, and input it to the decoder module.
[0073] The decoder module is used to restore the feature representation compressed by the encoder into a temperature sequence with the same dimension as the original input, thereby realizing the reconstruction process.
[0074] The goal of a deep autoencoder is to maximize the reconstruction accuracy of the original input. To achieve this goal, mean squared error (MSE) is used as the weight parameter of the autoencoder as the loss function. MSE quantifies the difference between the input x and the reconstructed output y, and its mathematical expression is shown in equation (4):
[0075]
[0076] Among them, L mse Let x represent the loss function constructed using mean squared error, x represent the input time-series thermal image sequence, and y represent the reconstruction approximation result generated by the decoder. n Let y represent the time-series thermal image sequence input at the nth time point. n This represents the approximate reconstruction result at the nth time point generated by the decoder, where N represents the time sequence length.
[0077] Building upon the autoencoder, multiple hidden layers are stacked to enable the model to learn more complex and hierarchical data patterns. For N thermal images of length H and width W, the temperature evolution of each pixel in the image is treated as new data. Through dimensionality transformation, the two-dimensional thermal image sequence is converted into W*H one-dimensional temporal data points of length N, which then serve as the input to the deep autoencoder.
[0078] Denoising Dataset: Although the latent space reduces noise interference by compressing information, the number of corresponding images is far less than that of the original images. Therefore, the output image reconstructed by the decoder is used as the final denoising result to construct a defect denoising dataset, which meets the data volume requirements of complex deep learning models. Specifically, based on the loss function in step (4), an L2 regularization term is added, and the mathematical expression is shown in equation (5):
[0079]
[0080] Among them, L reg Let w represent the L2 regularization loss function, λ represent the regularization coefficient, and w represent the loss function. m Let represent the m-th weight parameter in the decoder function, where M represents the total number of weight parameters. By penalizing excessively large weights in the network, overfitting is avoided, preventing the decoder from reconstructing noise information from the original input and reducing defect visibility.
[0081] S3, Step 3: Input the denoised time-series thermal image sequence into the Swin Transformer Wnet, and construct the total loss function to train the Swin Transformer Wnet to obtain the trained Swin Transformer Wnet. The Swin Transformer Wnet includes: extracting features from the denoised time-series thermal image sequence through the encoder Swin Unet to generate a latent spatial image, and reconstructing the generated latent spatial image into the original input time-series thermal image sequence through the decoder Swin Unet.
[0082] In this embodiment, a Swin Transformer Wnet model is trained using a constructed thermal image dataset. The network consists of two symmetrical Swin Unets, used to extract potential defect features from the spatial domain.
[0083] Specifically, the Swin Transformer Wnet consists of a cascaded encoder Swin Unet and a decoder Swin Unet forming a symmetrical structure. The encoder Swin Unet module is used to extract key features of defects and generate corresponding latent spatial images, which are then input to the decoder Swin Unet module. The decoder Swin Unet module is used to reconstruct the latent spatial images into the original input thermal images, ensuring that the pixel grayscale values of the two are as similar as possible.
[0084] Each Unet includes an image patch partitioning submodule, a linear embedding submodule, a Swing Transformer submodule, and an image patch merging submodule.
[0085] Image patch partitioning is used to divide the input image into non-overlapping 4*4 squares. Each square is flattened into a one-dimensional sequence that matches the input requirements of the Transformer model. Next, linear embedding is used to map each flattened image patch to a higher feature dimension through a fully connected layer, thereby extracting richer feature representations. Specifically, image patch partitioning and linear embedding are combined into a two-dimensional convolution operation, which improves computational efficiency and provides better input flexibility. The two-dimensional convolution operation is expressed as shown in Equation (6):
[0086] e = [Conv(x, D)] in D e kernel = ps, stride = ps)] T (6)
[0087] Where x is the original input and e is the output after two-dimensional convolution. T D represents the transpose operation. in and D e These represent the input and output channel dimensions of the two-dimensional convolution, respectively, while ps represents the size of the image patch, used to define the size of the convolution kernel and the stride.
[0088] Subsequently, the output of the two-dimensional convolution is input into the Swin Transformer module to capture complex nonlinear features in the image. The Swin Transformer consists of two consecutive modules, each consisting of layer normalization, a window-based multi-head self-attention layer, a residual connection layer, a normalization layer, and a multilayer perceptron. The window-based multi-head self-attention layer includes a window attention layer and a sliding window attention layer. The function expressions of different layers are shown in (7):
[0089]
[0090] In the formula, x l-1 This represents the input characteristics of the Swing Transformer submodule, corresponding to the output of layer l-1; This represents the output features of the window attention layer. x represents the output feature of the sliding window attention layer. l and x l+1 Let LN represent the l-th and (l+1)-th output features of the multilayer perceptron layer, Norm represent the normalization operation, WMSA and SWMSA represent the window attention operation and sliding window attention operation, and MLP represent the multilayer perceptron operation.
[0091] The normalization layer stabilizes the training process and accelerates the convergence speed. The window-based multi-head self-attention mechanism divides the feature map into non-overlapping blocks of equal size and performs self-attention within each window. The expression of the self-attention mechanism is shown in Equation (8):
[0092]
[0093] In the formula, Q, K, and / represent the query matrix, key matrix, and value matrix, respectively, and d k Let B represent the dimension of the key vector, and let B represent the relative positional bias, used to enhance the model's ability to handle long sequences. Window-based self-attention only occurs within local regions. A further constructed sliding window addresses the information isolation problem between windows, thus maintaining global information integration. The multilayer perceptron receives the normalized output of window self-attention, significantly enhancing the nonlinear representation capability of feature embeddings by expanding and compressing the channel dimension.
[0094] After passing through the Swin Transformer, the output data undergoes downsampling through image patch merging to achieve a hierarchical structure. First, the feature image is divided into non-overlapping 2×2 local windows. Then, pixels sharing the same spatial location in these windows are connected along the channel dimension to form a feature tensor with its width and height halved and its channel dimension multiplied by four. Finally, layer normalization is applied, and linear projection is performed to compress the channel dimension, resulting in a downsampling result with its width and height halved and its channel dimension multiplied by two. The mathematical expression of the above operations is shown in Equation (9):
[0095]
[0096] In the formula, i∈[0, H] l I2) and j∈[0, W l / 2), X 2i,2j X 2i,2j+1 X 2i+1,2j and This represents new features extracted from four different locations within a 2×2 image patch. Image patch merging utilizes a learnable linear layer to fuse features from adjacent pixels, thereby preserving local details and reducing information loss during downsampling.
[0097] Next, upsampling is performed to progressively restore resolution and detail. In the encoder, the Swin Unet, image patch expansion is introduced to achieve upsampling. Specifically, a linear mapping is first applied to double the number of feature channels. Then, the spatial resolution of the expanded features is doubled while the number of channels is reduced to one-quarter. Furthermore, skip connections are introduced to combine shallow encoder Swin Unet features with deep decoder Swin Unet features, thereby mitigating the degradation of spatial information during downsampling. Finally, a channel-wise linear transformation is performed using 1×1 convolutions, and the latent spatial image of the encoder output is obtained using the softmax activation function.
[0098] The structure of the decoder, Swin Unet, is similar to that of the encoder, with the main difference being that the linear projection layer of the decoder has a dimension of 256, corresponding to the grayscale range of an 8-bit image. Finally, the argmax activation function is applied to the linear projection output, producing a single-channel maximum probability index map as the final prediction output.
[0099] Further constructing the total loss function and training SwinTransformerWnet yields the trained SwinTransformerWnet, specifically:
[0100] Similar to deep autoencoders, Swin Transformer Wnet uses mean squared error as the reconstruction loss. Furthermore, since Swin Transformer Wnet can be viewed as a classification task, a cross-entropy loss function is introduced, its mathematical expression being shown in equation (10):
[0101]
[0102] In the formula, L ce Let y represent the cross-entropy loss function. c Let p represent the c-th real image. c This represents the predicted probability of the Swing TransformerWnet, where C represents the number of predicted classes.
[0103] Furthermore, a normalized cut loss from graph theory is introduced to constrain the latent spatial image generated by the encoder, and its mathematical expression is shown in Equation (11):
[0104]
[0105] In the formula, Ncut D (G) represents the normalized cut loss function, A d Let G represent the set of pixels belonging to category d, where D represents the number of categories, and G represents the set of pixels belonging to all categories in the image. `cut(A`d GA d The assoc(A) is used to calculate the similarity weights between pixels in category d and pixels in other categories. d G) is used to calculate the similarity weight between a pixel in category d and all pixels in all categories (including itself), (u, v) represents the pixel similarity weight, which measures the similarity between pixel u and pixel v, and w(u, t) represents the similarity weight between pixel u and pixel t.
[0106] The similarity weights of pixels u and v are calculated using the following formula:
[0107]
[0108] In the formula, I(u) represents the gray intensity of pixel u, I(v) represents the gray intensity of pixel v, ||I(u)-I(v)|| represents the gray intensity difference, and σ I The grayscale similarity scale parameter is represented by X(u), where X(v) represents the spatial coordinates of pixel u, X(v) represents the spatial coordinates of pixel v, and ||X(u)-X(v)|| represents the spatial Euclidean distance. σ X The parameter represents the spatial similarity scale, and r represents the neighborhood radius, used to limit the computational distance of the similarity weights. The normalized cut loss penalizes cross-subgraph connections by minimizing inter-class similarity and maximizing intra-class similarity, thereby achieving more accurate boundary segmentation.
[0109] In addition, a boundary openness loss based on morphological operations is introduced, as shown in Equation (13):
[0110]
[0111] In the formula, L open Let H denote the boundary open loss function, H represent the number of pixels in the feature image, P represent the latent space image generated by the encoder Swin Unet, p represent the pixel location, and I(p) represent the original gray value of pixel p. This indicates a morphological etching operation. This represents the morphological dilation operation, and S represents the structuring element.
[0112] Finally, by linearly combining all the cross-entropy loss functions, the normalized cut loss from graph theory, and the boundary opening loss based on morphological operations, a total loss function is constructed to balance their contributions, resulting in the trained Swing TransformerWnet.
[0113] Step 4: Input the new temporal thermal image sequence into the trained SwinTransformerWnet to generate the corresponding latent space image. Use a fully connected conditional random field for post-processing to obtain the final internal defect segmentation result.
[0114] In this embodiment, a fully connected conditional random field is applied to the corresponding latent space image generated by the Swing Transformer Wnet, and fine segmentation is achieved by minimizing the global energy function. The global energy function is shown in equation (14):
[0115]
[0116] In the formula, E(α) represents the global energy function of pixel label α, and ψ u Denotes a univariate potential function that causes the SwinTransformer Wnet to select labels with higher confidence, α i Let α represent the predicted label of the i-th pixel. j Let ψ represent the predicted label of the j-th pixel. p Let represent a binary potential function that penalizes similar pixels i and j by assigning them different labels.
[0117] To clearly demonstrate the unsupervised small-sample internal defect detection method based on Swin Transformer, experimental verification was conducted, such as... Figure 2 As shown, the first row contains the raw thermal images captured by an infrared camera, and the second row contains the detection results obtained by the unsupervised small-sample internal defect detection method based on Swin Transformer provided in this invention. The first three columns represent artificially created defects, and the last two columns represent real injection-molded bubble defects. The proposed Swin TransformerWnet model is trained on the circular acrylic defect data in the first column. After training, all model weight parameters are kept unchanged, and the model is directly applied to the detection of internal defects of different shapes and materials in columns 2 to 5. Figure 2 The results show that the model trained on a single material and defect shape can generalize well to different materials, shapes, and real injection molding defects without any additional training or fine-tuning. This indicates that the proposed method has good application prospects in practical industrial inspection, especially when defect samples are scarce and labeling is difficult.
[0118] In summary, this invention first constructs a deep autoencoder to compress the temporal thermal image sequence of defect samples into a low-dimensional latent space through unsupervised defect detection, improving the reconstruction accuracy of the original input image. A corresponding loss function is then constructed to preserve defect features, suppress thermal imaging noise, and adapt to the temporal correlation of thermal image sequences, effectively avoiding the limitations of traditional supervised defect detection. Furthermore, a Swing Transformer Wnet model is constructed, generating a latent space image through a continuous encoding and decoding structure. Detection is achieved with zero-labeled samples, and a corresponding loss function is constructed to optimize the image and boundary accuracy. Pixel-level optimization is achieved by combining fully connected conditional random field post-processing, ultimately outputting industrial-grade segmentation results. This enables accurate unsupervised internal defect detection under scarce sample conditions. Through this unsupervised architecture design, multimodal feature optimization, and industrial scenario adaptability, the three major technical bottlenecks of small sample size, high noise, and blurred boundaries in defect sample internal detection are solved, providing a high-precision, low-cost solution for sample internal defect detection in intelligent manufacturing.
[0119] Furthermore, the terms "upper," "lower," "inner," "outer," "front," and "rear" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Unless otherwise specifically stated, the relative steps, numerical expressions, and values of components and steps described in these embodiments do not limit the scope of the invention. Of course, the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of the invention. All equivalent changes or modifications made to the structures, features, and principles described in the claims of this invention should be included within the scope of the claims of this invention.
[0120] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An unsupervised small-sample internal defect detection method based on Swing Transformer, characterized in that, Includes the following steps: Step 1: Based on the pulsed thermal imaging system, the original infrared thermal image is acquired by adjusting the thermal imaging parameters of the sample to be tested, and after preprocessing, a time-series thermal image sequence of multiple defect samples is obtained. Step 2: Input the time-series thermal image sequence into the deep autoencoder, and construct the total loss function of the deep autoencoder to train the deep autoencoder to obtain the trained deep autoencoder. The deep autoencoder includes: compressing the input time-series thermal image sequence through the encoder to generate latent spatial features, and reconstructing the latent spatial features into a denoised time-series thermal image sequence through the decoder. Step 3: Input the denoised time-series thermal image sequence into the Swin Transformer Wnet, and construct the total loss function to train the Swin Transformer Wnet to obtain the trained Swin Transformer Wnet. The Swin Transformer Wnet includes: extracting features from the denoised time-series thermal image sequence through the encoder Swin Unet to generate a latent spatial image, and reconstructing the generated latent spatial image into the original input time-series thermal image sequence through the decoder Swin Unet. Step 4: Input the new temporal thermal image sequence into the trained Swing Transformer Wnet to generate the corresponding latent space image. Use a fully connected conditional random field for post-processing to obtain the final internal defect segmentation result.
2. The unsupervised small-sample internal defect detection method based on Swing Transformer according to claim 1, characterized in that, In step 1, the acquisition of the original infrared thermal image by adjusting the thermal imaging parameters of the sample to be tested includes: adjusting the heating time, defect depth and placement position of the sample to be tested in the pulse thermal imaging system to acquire the original infrared thermal image.
3. The unsupervised small-sample internal defect detection method based on Swing Transformer according to claim 1, characterized in that, In step 2, multiple hidden layers are stacked on top of the deep autoencoder. The encoder function and decoder function in the autoencoder of a single hidden layer are respectively... and Furthermore, both the encoder and decoder functions in the autoencoder use the hyperbolic tangent function tanh as the activation function, and the calculation formulas for the encoder and decoder functions are as follows: , , In the formula, This represents the input time-series thermal image sequence. This represents the latent space representation generated by the encoder. This represents the approximate reconstruction result generated by the decoder. and These represent the encoder's weight matrix and bias, respectively. and These represent the weight matrix and bias of the decoder, respectively.
4. The unsupervised small-sample internal defect detection method based on Swing Transformer according to claim 3, characterized in that, The loss function constructed in step 2 includes: using mean squared error to quantize the difference between the input and reconstructed output of the deep autoencoder, constraining the parameters of the deep autoencoder, and introducing an L2 regularization term to denoise the reconstructed approximation result. The calculation formula is as follows: , , , In the formula, Represents the total loss function. This represents the loss function constructed using the mean squared error. This represents the L2 regularization loss function. This represents the input time-series thermal image sequence. This represents the approximate reconstruction result generated by the decoder. Indicates the first A time-series thermal image sequence input at each time point, The decoder generates the first Approximate reconstruction results at each time point Indicates timing length, Represents the regularization coefficient. The first one in the decoder function One weight parameter, This represents the total number of weight parameters.
5. The unsupervised small-sample internal defect detection method based on Swing Transformer according to claim 4, characterized in that, In step 3, the encoder Swin Unet and the decoder Swin Unet are cascaded to form a symmetrical structure. Both the encoder Swin Unet and the decoder Swin Unet include an image block partitioning submodule, a linear embedding submodule, a SwinTransformer submodule, an image block merging submodule, and a skip connection submodule. The image block partitioning submodule is used to divide the denoised time-series thermal image sequence into multiple non-overlapping blocks, and flatten each block into a one-dimensional time-series thermal image sequence that matches the input requirements of the Swin Transformer submodule. The linear embedding submodule is used to map the flattened cube into high-dimensional features; The Swin Transformer submodule is used to generate a nonlinear feature image by taking the feature image output by two-dimensional convolution of multiple non-overlapping squares and high-dimensional features as input. The image block merging submodule is used to downsample the nonlinear feature image and then gradually restore the resolution and detailed information of the feature image through upsampling. The skip connection submodule is used to combine shallow features of the encoder Swin Unet with deep features of the decoder Swin Unet, perform linear transformations per feature channel using convolution, and obtain the latent spatial image of the encoder output using an activation function.
6. The unsupervised small-sample internal defect detection method based on Swing Transformer according to claim 5, characterized in that, The Swin Transformer submodule comprises two consecutive subunits. Each subunit includes a layer normalization layer, a window-based multi-head self-attention layer, a normalization layer, and a multilayer perceptron layer. The window-based multi-head self-attention layer includes a window-based multi-head self-attention layer and a sliding window-based multi-head self-attention layer. The function expressions for the different layers are as follows: , , , , In the formula, This represents the input characteristics of the Swing Transformer submodule, corresponding to Layer output; This represents the output features of the multi-head self-attention layer of the window. This represents the output features of the multi-head self-attention layer in the sliding window. and The first layer of a multilayer perceptron The and the first Output features: LN represents layer normalization operation, Norm represents normalization operation, WMSA represents multi-head self-attention operation of window and SWMSA represents multi-head self-attention operation of sliding window, and MLP represents multi-layer perceptron operation.
7. The unsupervised small-sample internal defect detection method based on Swing Transformer according to claim 5, characterized in that, Downsampling based on nonlinear feature images includes: The nonlinear feature image is divided into multiple non-overlapping local windows. Pixels sharing the same spatial position on the local windows are connected along the channel dimension to form a feature tensor with half the length and width and four times the channel dimension. Then, application layer normalization is applied, and linear projection is performed to compress the channel dimension, generating a downsampling result with half the length and width and twice the channel dimension.
8. The unsupervised small-sample internal defect detection method based on Swing Transformer according to claim 5, characterized in that, Then, the resolution and detailed information of the feature image are gradually restored through upsampling, including: Based on the downsampling results, a linear mapping is applied to double the number of feature channels, and the spatial resolution of the features is doubled. At the same time, the number of channels is reduced to one-quarter, gradually restoring the resolution and detailed information of the feature image.
9. The unsupervised small-sample internal defect detection method based on Swing Transformer according to claim 5, characterized in that, The total loss function of Swin Transformer Wnet is constructed by a linear combination of the cross-entropy loss function, the normalized cut loss in graph theory, and the boundary open loss based on morphological operations. The formula for calculating the cross-entropy loss function is as follows: , In the formula, Represents the cross-entropy loss function. Indicates the first A real image, This represents the predicted probability of the Swing Transformer Wnet. Indicates the number of predicted categories; The normalized cut loss in graph theory penalizes cross-subgraph connections by minimizing inter-class similarity and maximizing intra-class similarity, and the calculation formula is as follows: , In the formula, This represents the normalized cut loss function. Indicates belonging to a category The set of pixels, Indicates the number of categories. This represents the set of pixels of all categories in an image. Used to calculate categories The similarity weight between the middle pixel and other category pixels, Then used to calculate categories The similarity weight between the middle pixel and pixels of all categories. Represents the pixel similarity weight, measuring the pixel similarity. and pixels Similarity; pixel and pixels The formula for calculating the similarity weight is as follows: , In the formula, Represents pixels gray intensity Represents pixels gray intensity Indicates the difference in grayscale intensity. This represents the grayscale similarity scale parameter. Represents pixels spatial coordinates, Represents pixels spatial coordinates, Represents spatial Euclidean distance. Represents the scale parameter of spatial similarity. Represents the neighborhood radius, used to limit the distance for calculating similarity weights; The formula for calculating the boundary openness loss based on morphological operations is as follows: , In the formula, This represents the boundary open loss function. This indicates the number of pixels in the feature image. This represents the latent spatial image generated by the encoder SwinUnet. Indicates the position of a pixel. Represents pixels The original grayscale value, This indicates a morphological etching operation. This indicates a morphological dilation operation. Represents a structural element.
10. The unsupervised small-sample internal defect detection method based on Swing Transformer according to claim 5, characterized in that, In step 4, a fully connected conditional random field post-processing method is used to obtain the final internal defect segmentation result, including: for the corresponding latent space image generated by the Swin Transformer Wnet, a fully connected conditional random field is applied to achieve fine segmentation by minimizing the global energy function. The calculation formula is as follows: , In the formula, Represents pixel label The global energy function, This represents a univariate potential function that allows Swin TransformerWnet to select labels with higher confidence. Indicates the first Predicted label for each pixel Indicates the first Predicted label for each pixel Represents a binary potential function, penalizing similar pixels. and Assign different labels.
Citation Information
Patent Citations
Defect detection method based on Swin Transformer structure in combination with SSIM and GMSD
CN114841977A
Industrial assembly scene low data dependence defect detection method based on Swin Transform model
CN118797469A