Visual defect detection method based on multi-level Transform

By combining a multi-layered Transformer architecture with a pre-trained convolutional network, the problems of multi-scale feature fusion and global-local dependency modeling in unsupervised defect detection are solved, achieving efficient defect detection that is suitable for industrial scenarios where defect samples are scarce.

CN121329967AActive Publication Date: 2026-01-13SOUTHEAST UNIV

Patent Information

Application Number
CN202511808902.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-01-13
Estimated Expiration
2045-12-03

AI Technical Summary

Technical Problem

Existing unsupervised defect detection methods struggle to effectively integrate multi-scale features, fail to fully model global and local dependencies, and lack detection accuracy when defect samples are scarce.

Method used

We design a visual defect detection method based on multi-level Transformer, which combines multi-scale feature extraction of pre-trained convolutional networks with the global modeling capability of Transformer. Through a multi-level encoder-decoder architecture, we achieve multi-scale feature fusion and anomaly scoring. The unsupervised learning paradigm requires only normal samples for training.

Benefits of technology

It achieves efficient detection of complex textures and minute defects, reduces model complexity and computational overhead, is suitable for industrial scenarios where defect samples are scarce, and improves detection accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121329967A_ABST
    Figure CN121329967A_ABST
Patent Text Reader

Abstract

The invention discloses a visual defect detection method based on multilevel Transform. The method comprises the following steps: firstly, extracting multi-scale features of an input image through a pre-trained convolutional neural network, and performing hierarchical fusion; then, a multi-level Transform coding and decoding structure is used for carrying out deep reconstruction on the fusion features, and long-distance dependency relationships of different granularities are captured through hierarchical patch segmentation and a self-attention mechanism; meanwhile, introducing a standard deviation to estimate branch learning reconstruction uncertainty; and finally, a pixel-level anomaly score graph is generated by calculating a normalized residual error and adopting a multi-scale anomaly score aggregation strategy, so that accurate anomaly positioning and discrimination are realized. According to the method, multi-scale feature representation and global semantic modeling capability are fused, the accuracy, robustness and generalization capability of anomaly detection are effectively improved, only normal sample training is needed, and the method is suitable for complex scenes such as industrial visual detection and has a good application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary fields of computer vision, deep learning and industrial defect detection, and specifically relates to a visual defect detection method based on multi-level Transformer. Background Technology

[0002] In modern industrial production, product quality control is a crucial link in ensuring production efficiency and product competitiveness. Defect detection, as a core technology of quality control, is of great significance for timely detection of product defects, reducing scrap rates, and improving production efficiency. Traditional defect detection methods mainly rely on manual visual inspection or image processing techniques based on manual features. These methods are effective for simple and regular defects, but as industrial products become increasingly complex and diverse, the limitations of traditional methods are becoming increasingly apparent.

[0003] In recent years, the rapid development of deep learning technology has brought new solutions for defect detection. Supervised learning methods based on convolutional neural networks (CNNs) can achieve high detection accuracy under conditions of sufficient labeled samples. However, supervised learning methods face the following key problems: First, in real-world industrial scenarios, defect samples are often scarce and difficult to obtain, making the construction of large-scale, high-quality labeled datasets extremely costly; second, defect types are numerous and unevenly distributed, making it difficult to exhaustively represent all possible defect patterns; third, traditional convolutional networks mainly focus on local feature extraction and lack the ability to model long-distance dependencies in images, easily leading to incomplete reconstruction or insufficient feature representation.

[0004] To address the aforementioned issues, unsupervised defect detection methods have emerged. Unsupervised methods require only training with normal samples, learning the data distribution of normal samples, and detecting anomalous samples that deviate from the normal distribution during testing. Existing unsupervised defect detection methods primarily employ autoencoders or generative adversarial networks (GANs) for image reconstruction, identifying defects by comparing the differences between the original and reconstructed images. However, these methods mostly employ single-scale, shallow reconstruction strategies, making it difficult to simultaneously capture both global semantic information and local detail features of the image, resulting in limited detection capabilities for complex textures or minute defects.

[0005] In recent years, the Transformer architecture has achieved great success in the field of natural language processing, with its self-attention mechanism effectively modeling long-distance dependencies in sequences. Subsequently, the Vision Transformer (ViT) applied the Transformer to computer vision tasks, demonstrating powerful feature representation capabilities. However, directly applying the Transformer to defect detection still faces challenges: on the one hand, standard Transformers require a large amount of data for training, while the number of normal samples in defect detection scenarios is limited; on the other hand, single-layer Transformers struggle to process feature information at different scales simultaneously, failing to fully utilize the important role of multi-scale features in defect detection.

[0006] Furthermore, existing methods also have shortcomings in feature fusion strategies. Simple feature concatenation or addition is insufficient to fully exploit the complementary information between features at different levels, while overly complex fusion networks increase the training difficulty and computational cost of the model. How to design an efficient multi-scale feature fusion and reconstruction mechanism that controls model complexity while ensuring detection accuracy is a problem that current technology urgently needs to solve.

[0007] Therefore, there is an urgent need for an unsupervised defect detection method that can effectively integrate multi-scale features, fully model global and local dependencies, and is applicable to limited training samples. Summary of the Invention

[0008] To address the aforementioned issues, this invention discloses a visual defect detection method based on a multi-layered Transformer. It innovatively designs a multi-layered encoder-decoder architecture, combining the multi-scale feature extraction capabilities of a pre-trained convolutional network with the global modeling capabilities of the Transformer, achieving high-quality learning and reconstruction of normal sample patterns. In the detection phase, by comparing the differences between the reconstructed features and the original features, and combining this with a multi-scale anomaly scoring mechanism, it achieves accurate localization and discrimination of various defects.

[0009] To achieve the above objectives, the technical solution of the present invention is as follows:

[0010] A visual defect detection method based on multi-level Transformer includes five stages: feature extraction, multi-level encoding, multi-level decoding, reconstruction and prediction, and defect discrimination. The specific steps are as follows:

[0011] S1: Feature Extraction Stage

[0012] The input image is processed through a pre-trained convolutional neural network backbone to extract multi-scale features, and the features from different levels are fused to obtain multi-scale fused features.

[0013] Specifically, a pre-trained ResNet18 network is used as the feature extractor, extracting feature maps from three different layers (layer 1, layer 2, and layer 3) through a forward propagation hook function. A feature embedding concatenation function is then used to fuse the multi-scale features: for two feature maps... and ,in Spatial resolution Higher than Spatial resolution First, calculate the step size. ,right Perform an unfold operation to divide it into blocks, and then correlate it with the channel dimension. The features are then concatenated, and finally the spatial structure is restored using a fold operation. This embedding concatenation operation is recursively applied to fuse the three layers of features layer by layer, resulting in a fused feature representation containing multi-scale semantic information. .

[0014] S2: Multi-level coding stage

[0015] Multi-scale fused features are input into a multi-level Transformer encoder. The features are then downsampled and semantically extracted at multiple levels through a hierarchical coding structure to obtain latent representations of different granularities.

[0016] Specifically, firstly, the fusion features are... Convolutional mapping to the input dimension of Transformer Apply global positional encoding to the feature map. This location encoding uses a learnable embedding method to add location information to each spatial location in the feature map:

[0017]

[0018] in Embed parameters for learnable locations.

[0019] Flatten the feature map into a sequence form, and transform the dimensionality to... ,in For batch size, and The height and width of the feature map, This represents the number of channels.

[0020] For each coding level Perform the following operations:

[0021] (1) Arrange the input features according to Perform block reorganization;

[0022] (2) Add a learnable zero vector as a global token before each patch;

[0023] (3) Applying local location coding Add relative position information to each position within the patch;

[0024] (4) Self-attention calculation and feedforward network processing are performed using Transformer Encoder;

[0025] (5) Separate the global token and the local token. The global token is passed to the next layer as an abstract representation of this layer, and the local token passes through the bottleneck layer. The compressed data is saved for the decoding stage.

[0026] The TransformerEncoder includes a multi-head self-attention mechanism and a feedforward neural network. The calculation formula for the multi-head self-attention mechanism is as follows:

[0027]

[0028] in, , , These are query, key, and value matrices, respectively. This is the transpose of the matrix. The dimension of the key vector. For the number of attention heads, This is for outputting the projection matrix.

[0029] The feedforward neural network uses two fully connected layers with an activation function in between:

[0030]

[0031] in For GELU or ReLU activation functions, , This is the weight matrix. , This is the bias vector.

[0032] Bottleneck layer The following structure is used, consisting of two linear layers and a layer normalization layer, with the GELU activation function used in between:

[0033]

[0034] in From the dimension Compress to Restore the dimension to This structure is used to compress local features of the encoder output, reducing storage overhead, while retaining key semantic information for feature reconstruction during the decoding stage.

[0035] S3: Multi-level decoding stage

[0036] The latent representation is upsampled and reconstructed layer by layer through a multi-level Transformer decoder to restore the spatial resolution of the original features layer by layer.

[0037] Specifically, decoding is performed layer by layer upwards, starting from the deepest level. For each decoding level... Perform the following operations:

[0038] (1) The global token of this level and the local token saved during the encoding phase The sequences are concatenated along the sequence dimension and used as the memory input for the decoder.

[0039] (2) Use learnable query embeddings As the target input of the decoder The number of query embeddings is equal to ;

[0040] (3) Cross-attention calculation is performed through TransformerDecoder to make the query vector focus on the features of the encoding stage;

[0041] (4) Reassemble the decoded output into a spatial structure to restore a higher resolution feature map;

[0042] Repeat the above operation, upsampling layer by layer until the resolution of the original feature map is restored. The final decoded output feature dimension is... .

[0043] The computation process of TransformerDecoder includes: self-attention calculation of the target sequence, enabling the query vectors to interact with each other; encoder-decoder cross-attention calculation, enabling the query vectors to focus on the output features of the encoder; and feedforward neural network processing.

[0044] In the cross-attention mechanism, From decoder input, and The calculation formula is derived from the encoder output:

[0045]

[0046] in This is the query matrix for the decoder. and These are the key matrix and value matrix output by the encoder, respectively.

[0047] S4: Restructuring and Prediction Phase

[0048] The features output by the decoder are mapped back to the original feature space through a convolutional layer to obtain reconstructed features and anomaly score maps.

[0049] Specifically, the features output by the decoder are from Dimensional transformation back Format. (Through the first...) Convolutional layer The features are mapped back to the number of channels of the original fused features to obtain the reconstructed features. .

[0050] The reconstructed features are detached to block gradient propagation, and then passed through a second multi-layer convolutional network. Processing:

[0051] (1) Passing through multiple in sequence Convolution, batch normalization, and LeakyReLU activation function

[0052] (2) Gradually increase the number of channels from Dimensional reduction , ,

[0053] (3) Finally passed Anomaly score map of single channel output by convolution

[0054] Return to reconstructed features and anomaly score plot As model output.

[0055] S5: Defect Identification Stage

[0056] The difference between the reconstructed features and the original features is calculated, and the defect is located and identified by combining the anomaly score map.

[0057] Specifically, compute the reconstructed features Features of original fusion Between Norm distance :

[0058]

[0059] Norm calculation is performed on the channel dimension while preserving the spatial dimension.

[0060] Perform multi-scale pooling on the normalized distance map: using They are respectively , , , Average pooling; upsample the pooling results at each scale back to the original resolution; concatenate the features at all scales.

[0061] For the For each scale, the pooling operation formula is:

[0062]

[0063] in .

[0064] Weighted fusion of the spliced ​​multi-scale features:

[0065]

[0066] Among them, weight .

[0067] The final anomaly score map is upsampled to the input image resolution and then thresholded to obtain a binary mask of the defect region.

[0068] The model training process includes:

[0069] Training is performed using only defect-free, normal samples. The loss function is defined as the mean squared error loss.

[0070]

[0071] in, To reconstruct the loss, Predict loss for outlier scores. This indicates the expected operation.

[0072] Use the Adam optimizer to update parameters, with the learning rate set to... During training, the backbone network parameters are kept fixed and not updated; only the parameters of the Transformer encoder, decoder, and convolutional layers are trained. When the loss on the validation set no longer decreases, the model checkpoint is saved.

[0073] The present invention has the following beneficial effects:

[0074] 1. Powerful multi-scale feature fusion capability: Through an innovative feature embedding and concatenation mechanism, this invention can effectively fuse features from different layers of the pre-trained network, preserving low-level detail information and high-level semantic information while avoiding feature redundancy. Compared to simple feature concatenation or addition, this method achieves deep interaction of features through unfold-fold operations, significantly improving feature representation capabilities.

[0075] 2. Efficient Global-Local Dependency Modeling: The multi-layered Transformer architecture achieves progressive modeling from global semantics to local details through the separate processing of global and local tokens. The self-attention mechanism can capture the dependencies between any locations in the image, overcoming the problem of limited receptive field in traditional convolutional networks and significantly improving the detection capability of complex textures and minute defects.

[0076] 3. Flexible hierarchical encoding / decoding strategy: By employing different patch sizes at different levels, this method achieves progressive feature abstraction from fine-grained to coarse-grained, and layer-by-layer reconstruction from abstract to concrete. This hierarchical strategy enables the model to handle defects of different scales simultaneously, improving detection robustness and generalization ability.

[0077] 4. Unsupervised learning paradigm, reducing data dependence: This invention adopts an unsupervised learning paradigm, requiring only normal samples for training, eliminating the need for expensive defect sample annotation. This significantly reduces deployment costs in real-world industrial scenarios, making it particularly suitable for applications where defect samples are scarce or defect types are diverse.

[0078] 5. Precise multi-scale anomaly scoring mechanism: In the defect discrimination stage, multi-scale pooling and weighted fusion strategies are used to integrate anomaly signals from different receptive fields, improving the detection sensitivity for defects of different sizes. The normalized distance map combined with the predicted anomaly score map further enhances the discriminative power of anomaly regions.

[0079] 6. Excellent computational efficiency and scalability: By compressing local features through a bottleneck layer, storage overhead and computational complexity are effectively reduced. The backbone network parameters are fixed and not updated; only the Transformer part is trained, accelerating the training process. The model structure is clear and easy to deploy and optimize on different hardware platforms.

[0080] 7. Broad application prospects: This method has achieved excellent detection performance on public datasets such as MVTec AD. It performs well in defect detection tasks of various materials such as fabrics, metals, plastics, and wood, and has good generalization ability and practical application value. Attached Figure Description

[0081] Figure 1This is a flowchart of the overall process of the method of the present invention, which shows the complete processing flow from the input image to the defect discrimination result.

[0082] Figure 2 This is a diagram of the multi-layer Transformer encoder structure, illustrating the specific implementation of layered encoding, global / local token separation, and bottleneck layer compression. Detailed Implementation

[0083] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0084] The core innovation of this invention lies in:

[0085] 1. Multi-scale feature embedding and concatenation mechanism: A recursive feature embedding and concatenation method is proposed. Through the unfold-concat-fold operation, feature maps of different levels of the pre-trained network are deeply fused. While preserving semantic information at each scale, it avoids the feature redundancy problem caused by simple concatenation.

[0086] 2. Multi-layered Transformer encoding and decoding architecture: A layered downsampling and upsampling Transformer structure was designed. At each layer, through patching and global token mechanisms, progressive feature abstraction from fine-grained to coarse-grained and layer-by-layer reconstruction from abstract representation to specific features were achieved.

[0087] 3. Dual-position encoding strategy: Introducing a combination of global position encoding and local position encoding, global encoding captures the overall spatial layout of the feature map, while local encoding models the relative positional relationships within the patch, enhancing the model's ability to perceive spatial structure.

[0088] 4. Bottleneck layer feature compression: Local tokens are compressed through the bottleneck layer during the encoding stage, and the compressed features are used as memory input during the decoding stage, which effectively reduces storage overhead while preserving key semantic information.

[0089] 5. Multi-scale anomaly scoring fusion: In the defect discrimination stage, a multi-scale pooling and weighted fusion strategy is adopted to integrate anomaly signals under different receptive fields, thereby improving the robustness of detection for defects of different sizes.

[0090] Specific examples Figure 1 As shown, the visual defect detection method based on multi-level Transformer of this invention consists of five stages: feature extraction, multi-level encoding, multi-level decoding, reconstruction and prediction, and defect discrimination. The specific implementation methods of each stage are described in detail below.

[0091] S1: Extract multi-scale features from the input image using a pre-trained convolutional neural network backbone, and fuse features from different levels to obtain multi-scale fused features. The specific steps are as follows:

[0092] S1.1: Input the original image into the feature extraction module;

[0093] Specifically, the original images are divided into defective and defect-free images. During the training phase, only defect-free images are used, allowing the model to learn the data distribution of defect-free samples. During the inference phase, the use of defective images is not restricted; the model will reconstruct the input image and output the discrimination result. The training images undergo preprocessing: first, the images are adjusted to... Pixel (in this embodiment) Then, bicubic interpolation is applied for scaling to... Pixels, apply padding (padding size is...) (pixels), random rotation Degree, randomly cropped to Pixels, converted to tensor format, and using the mean of the ImageNet dataset. and standard deviation Normalize.

[0094] S1.2: Use a pre-trained ResNet18 network as the backbone feature extractor;

[0095] Specifically, ResNet18 is a convolutional neural network pre-trained on the ImageNet dataset, containing four main layers (layer 1 to layer 4). In this embodiment, intermediate feature maps from three different layers (layer 1, layer 2, and layer 3) are extracted using a forward propagation hook function. The hook function is implemented as follows: a list of outputs is defined to store features, a hook function is defined to add the output features to the list, and then the `register_forward_hook` method is used to register the hook function to the output of the specified layer. For an input size of... The images from which the three feature maps were extracted have the following dimensions:

[0096] (1) Layer 1 output: Where B is the batch size

[0097] (2) Layer 2 output:

[0098] (3) Layer 3 output:

[0099] In this embodiment, layer 4 features are not used to achieve a balance between feature representation and computational efficiency. The backbone network parameters remain fixed during training and no gradient updates are performed, thereby making full use of pre-trained knowledge and reducing training time.

[0100] S1.3: Multi-scale features are fused using the feature embedding concatenation function (embedding_concat);

[0101] Specifically, for two feature maps with different spatial resolutions and ,in The spatial resolution is The number of channels is , The spatial resolution is The number of channels is The feature embedding and concatenation process is as follows:

[0102] (1) Calculate the step size ;

[0103] (2) For high-resolution feature maps Perform the unfold operation: using Parameters, will Decompose into sizes of Non-overlapping patches, output dimension is The third dimension represents the content within each patch. Each pixel is expanded into an independent channel;

[0104] (3) Initialize the zero tensor , dimension ;

[0105] (4) To Loop through each position and concatenate them along the channel dimension:

[0106] ,in from arrive ;

[0107] (5) Reshape the spliced ​​tensor as follows: Two-dimensional form;

[0108] (6) Use the fold operation to restore the spatial structure: Parameters, reorganizing features into Spatial form.

[0109] The mathematical representation of this process is:

[0110]

[0111] Compared to simple feature splicing or addition, this embedding and splicing method can achieve deep interaction of features at different scales while preserving spatial information at multiple scales.

[0112] S1.4: Recursively apply the embedding and splicing operation to fuse the three-layer features layer by layer to obtain a fused feature representation containing multi-scale semantic information.

[0113] Specifically, the extracted three-layer features are embedded and concatenated twice:

[0114] (1) Fusing features from layer 1 and layer 2

[0115] in Features of layer 1 , For layer 2 features After fusion, the dimension of outputs_temp is

[0116] (2) Fuse the fusion result with the layer3 features again.

[0117] in Layer 3 features The final fused feature outputs_final has a dimension of

[0118] In this embodiment, the fused feature dimension is Channel, spatial resolution is This fusion feature simultaneously includes shallow texture detail information (from layer 1), mid-level local pattern information (from layer 2), and deep semantic abstraction information (from layer 3), providing rich multi-scale representations for subsequent multi-level Transformer encoding.

[0119] S2: Input the multi-scale fused features into a multi-level Transformer encoder. Through a hierarchical coding structure, perform multi-level downsampling and semantic extraction on the features to obtain latent representations of different granularities. The specific steps are as follows:

[0120] S2.1: Map the fused features to the input dimension of the Transformer through a convolutional layer;

[0121] Specifically, use one The convolutional layer (pre_conv) fuses features from Channel mapping to Channels. This convolutional layer does not change the spatial resolution, but only transforms the channel dimension. The mapped feature dimension is... In this embodiment, a convolutional layer without bias (bias=False) is used to reduce the number of parameters.

[0122] S2.2: Apply global positional encoding to the feature map;

[0123] Specifically, global position encoding employs a learnable embedding approach, implemented through the `LearnedPositionEncoding1` class. This class inherits from `nn.Embedding` and maintains a learnable position embedding matrix with parameters of shape [ ,Right now During forward propagation, this matrix is ​​reshaped as follows: And add the batch dimension, that is Then, add it element-wise to the input features:

[0124]

[0125] Here, `position_weight` is a learnable parameter that automatically learns global positional information for each spatial location in the feature map through training. After adding positional encoding, a Dropout layer is applied. Regularization is applied. Global location encoding enables the model to distinguish information from different spatial locations in the feature map, which is crucial for capturing the global structure of the image.

[0126] S2.3: Flatten the feature map into a sequence form;

[0127] Specifically, the two-dimensional spatial feature map Convert to sequence form The transformation process is as follows: First, the features are permute... Then the view is Finally, unsqueeze(0) increases the sequence dimension. For this embodiment... Features, after conversion This representation treats each spatial location as a token, and the entire feature map becomes a string of length 1. The sequence conforms to the input format requirements of Transformer.

[0128] S2.4: For each coding level, perform the following operations:

[0129] This embodiment sets up 4 encoding levels, with patch sizes (patch_size) as follows: The following is the first... The encoding process is illustrated using layers as an example, and this process is implemented by the forwardDOWN method:

[0130] (1) Feature block recombination:

[0131] First, calculate the patch size for the current level. And the cumulative number of patches in the lower level For the first layer, , For example, layer 0 ( )hour, .

[0132] Then the input sequence (dimension) The features are reorganized into a patch structure using view and permute operations. The form in which the first dimension The second dimension represents the number of tokens in the current patch. This indicates the number of patches within a batch.

[0133] (2) Add a learnable global token:

[0134] A zero vector is added at the beginning of each patch sequence as a global token (class token). The zero vector has dimensions of... Expanded to This matches all patches within a batch. The global token aggregates information from the entire patch, similar to the CLS token in Vision Transformer.

[0135] (3) Apply local location coding:

[0136] Use LearnedPositionEncoding2 to add relative position information to each token within the patch. The position encoding dimension is... ,pass Expand to This is then added to the local token. Local positional encoding enables the model to learn the spatial structural relationships within the patch.

[0137] (4) Processed by TransformerEncoder:

[0138] The token sequence that combines the global token and the token with local location encoding. Input TransformerEncoder. TransformerEncoder contains... Layers (each encoding level in this embodiment contains 4 layers), each layer includes:

[0139] a) Multi-Head Self-Attention Mechanism:

[0140] Use 8 attention points ( Each head has a dimension of The calculation process is as follows: First, a query is generated through a linear transformation. ,key ,value The matrix is ​​then used to calculate the attention weights.

[0141]

[0142] Finally, perform a weighted sum:

[0143]

[0144] b) Feed-Forward Network:

[0145] It employs two fully connected layers, with the middle layer having a dimension of [missing information]. The activation function uses GELU (in this embodiment, the activation parameter is set to "gelu"). The calculation process is as follows:

[0146]

[0147] c) Residual connections and layer normalization: Add residual connections and LayerNorm after each sublayer to enhance training stability.

[0148] (5) Separate the global token and the local token:

[0149] In the sequence output by the encoder, the first token (index) ) is the global token, denoted as latent_patch, dimension This global token aggregates global information for the current patch and will be passed as an abstract representation to the next layer of encoding.

[0150] The remaining Each token is a local token, denoted as latent_pixel, with dimensions... These local tokens retain detailed spatial information within the patch.

[0151] (6) Compress local tokens through the bottleneck layer:

[0152] The local token `latent_pixel` is compressed using a bottleneck layer. The bottleneck layer consists of two linear layers and a layer normalization mechanism, with the following structure:

[0153] The above is a specific implementation example of the bottleneck structure in the invention. This structure first changes the feature dimension from... Compress to After activation, it will be restored to its original state. This "compression-recovery" design preserves key semantic information while also acting as a regularization mechanism to prevent overfitting. The compressed local tokens are stored in `latent_list` for use in the decoding phase.

[0154] S2.5: Repeat step S2.4 to complete the processing of all coding levels and obtain multi-level potential representations.

[0155] Specifically, the above encoding operation is performed iteratively across the four encoding levels. The input to each level is the global token (latent_patch) of the previous level, and the output is the global token and local token of the current level. The processing dimensions change as follows: Level 0 Input is After reorganization, it becomes ( A local token, (one patch), after adding the global token is... The final output global token is The output local token (after compression) is: ;No. layer Input is After reorganization, it becomes Output the global token as The output local token (after compression) is: ;No. layer Input is After reorganization, it becomes The output global token is [ The output local token (after compression) is: ; third floor Input is After reorganization, it becomes Output the global token as The output local token (after compression) is: .

[0156] Through this hierarchical encoding, features are derived from fine-grained ( The spatial resolution is gradually abstracted to a coarse-grained level (ultimately, each image is represented by only one global token). The global token at each level captures the abstract representation at that scale, while the local tokens preserve the spatial details at that scale. Together, they constitute a multi-layered latent representation.

[0157] S3: The latent representation is upsampled and reconstructed layer by layer using a multi-level Transformer decoder to restore the spatial resolution of the original features. The specific steps are as follows:

[0158] S3.1: Start decoding from the deepest level and proceed upwards layer by layer;

[0159] Specifically, the decoding process is the reverse of the encoding process, starting from layer 3 (the deepest layer) and proceeding layer by layer towards layer 0 (the shallowest layer). The decoding order is as follows: The purpose of each decoding layer is to progressively restore the abstract global representation to specific spatial features, until it is restored to the original representation. Resolution.

[0160] S3.2: For each decoding level, perform the following operations:

[0161] The following is the first The decoding process is illustrated using a layer as an example, and this process is implemented by the forwardUP method:

[0162] (1) Concatenate the global token and the local token:

[0163] The global token (latent_patch, dimension) saved during the encoding phase of this level ) and compressed local token (latent_pixel, dimension) The sequence is concatenated along its dimension to obtain the complete memory input, with a dimension of [dimension number missing]. The splicing method is implemented as follows:

[0164]

[0165] (2) Use learnable query embeddings:

[0166] Define a learnable query embedding, which is an nn.Embedding layer with parameters of shape [formula missing]. During decoding, the query is embedded. Expand to Then expand to To match all patches. The query embedding represents the target pattern that this level wants to reconstruct, and learns the optimal query method through training.

[0167] (3) Perform cross-attention calculation using TransformerDecoder:

[0168] embedding queries as Input, concatenation As memory input, it is processed by TransformerDecoder. TransformerDecoder contains... Layers (4 layers in this embodiment), each layer includes three sub-modules:

[0169] a) Self-Attention of the Target Sequence:

[0170] Self-attention is computed between query embeddings, enabling interaction between different query positions. The computation process is the same as that of the encoder, but the input is only the query sequence.

[0171] b) Encoder-decoder cross-attention:

[0172] This is the core mechanism of the decoder, enabling the query vector to focus on features from the encoding stage. Query input from the decoder, and The memory output from the encoder. Cross-attention is calculated as follows:

[0173]

[0174] Through this mechanism, the decoder can selectively extract the information from the encoder that is most relevant to the current reconstruction target.

[0175] c) Feedforward Neural Network (FFN):

[0176] Similar to the FFN structure in the encoder, it contains two fully connected layers and a GELU activation function. Residual connections and layer normalization are added after each submodule. After processing by multiple decoders, the output sequence has a dimension of... This sequence contains all the information needed for reconstruction.

[0177] (4) Reassemble the decoded output into a spatial structure:

[0178] The sequence output by the decoder It is necessary to reconstruct the spatial features at a higher resolution. The reconstruction process is as follows:

[0179] First, view the sequence as... The first dimension is about to be completed. Split into two spatial dimensions, with the second dimension It is split into batch and two patch dimensions.

[0180] Then through Operation, rearrange the dimensions as This ensures that batch, patch, and spatial dimensions correspond correctly.

[0181] Finally passed The operation flattens the features into a sequence with dimension . .

[0182] This step achieves spatial resolution recovery from coarse-grained to fine-grained. For example, if Then the spatial resolution starts from Restore to .

[0183] S3.3: Repeat step S3.2, upsampling layer by layer until the resolution of the original feature map is restored;

[0184] Specifically, the above decoding operation is performed iteratively across the four decoding levels. The processing dimensions of each layer change as follows: Layer 3 decoding Enter the global token as The input local token is splicing memory as Query embedding is The output is Layer 2 decoding Input is memory is Query embedding is The output is Layer 1 Decoding Input is memory is Query embedding is The output is Layer 0 decoding Input is memory is Query embedding is The output is .

[0185] S3.4: The feature dimension of the final decoded output is .

[0186] Specifically, after all decoding levels, the final output dimension is ,in For spatial resolution, This represents the number of feature channels. The sequence is then processed... Remove the first dimension, then pass through Reorganized into Finally passed Adjust the dimension order as follows The features are then restored to the standard feature map format. At this point, the decoding process is complete, and the features are gradually restored from the most abstract global representation to the same spatial resolution as the input fused features.

[0187] S4: Map the features output by the decoder back to the original feature space through a convolutional layer to obtain the reconstructed features and anomaly score map.

[0188] Specifically, two independent convolutional networks are used to generate reconstructed features and anomaly score maps, respectively:

[0189] (1) Reconstruction feature generation: via final_layer1 (a The convolutional layer will output the decoder. Mapping back to the number of channels in the original fused features yields the reconstructed features. This convolutional layer has no bias (bias=False).

[0190] (2) Anomaly score map generation: After performing a detach() operation on the reconstructed features to block gradient propagation, the features are processed through final_layer2 (a multi-layer convolutional network). This network contains... The convolutional layer has the following structure:

[0191] The first layer is ,

[0192] The second layer is ,

[0193] The third layer is ,

[0194] The 4th layer is ,

[0195] The 5th floor is All convolutional layers Gradually increase the number of channels from Dimensional reduction The final output is a single-channel anomaly score graph. , representing the standard deviation estimate (uncertainty) for each location.

[0196] S5: Calculate the difference between the reconstructed features and the original features, and combine the anomaly score map to locate and identify defects.

[0197] Specifically, the defect identification process includes the following steps:

[0198] (1) Calculate the normalized distance: calculate the distance between the reconstructed features and the original fused features. Normal distance, divided by the absolute value of the outlier score plot for normalization:

[0199]

[0200] in, Norm calculation in the channel dimension ( Perform on, keep To preserve spatial dimensions. The normalized distance map has the following dimensions: .

[0201] (2) Multi-scale pooling: Perform multi-scale average pooling on the normalized distance map, using... Different patch_sizes: , , , For the first One scale, The pooling operation for each scale is as follows:

[0202]

[0203] The pooled features are upsampled back to the original resolution using bilinear interpolation. ,get Anomaly score plots at different scales.

[0204] (3) Feature splicing and weighted fusion: Initialize a zero tensor The embedding_concat function is used to concatenate the features at the four scales sequentially to obtain... Multi-scale features. Then through Convolutions are weighted and fused, with convolution weights being... The fusion formula is:

[0205]

[0206] (4) Final anomaly score calculation: Upsample the fused anomaly score map to the input image resolution (in this embodiment, 10 ... Using bilinear interpolation mode (mode="bilinear"), the anomaly score plots for all test samples are normalized.

[0207]

[0208] Image-level anomaly score is the maximum anomaly score for each image. Pixel-level discrimination generates a binary mask by setting a threshold on the normalized score image, enabling precise localization of defect areas.

[0209] In model training, this embodiment uses the Adam optimizer, with a learning rate set to... .train epochs, batch size is ,use Each CPU thread loads data. The loss function consists of two parts:

[0210] (1) Reconstruction loss: It measures the difference between the reconstructed features and the original features;

[0211] (2) Standard deviation prediction loss: This allows the anomaly score map to learn and predict the distribution of reconstruction errors.

[0212] The total loss is the sum of the two. The parameters of the backbone network (ResNet18) are fixed and not updated; only the parameters of the Transformer encoder / decoder, bottleneck layer, and convolutional layer are trained. When the loss on the validation set no longer decreases, the model checkpoint is saved. On the MVTec AD dataset, this method achieves excellent performance on both image-level and pixel-level detection tasks, validating the effectiveness of the proposed method.

[0213] It should be noted that the above content merely illustrates the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. For those skilled in the art, various improvements and modifications can be made without departing from the principle of the present invention, and all such improvements and modifications fall within the scope of protection of the claims of the present invention.

Claims

1. A method for visual defect detection based on multi-level Transformer, characterized in that: The specific steps are as follows: S1: feature extraction stage The input image is extracted through a pre-trained convolutional neural network backbone to obtain multi-scale features, and the features of different levels are fused to obtain multi-scale fusion features; S2: multi-level encoding stage The multi-scale fusion features are input into a multi-level Transformer encoder to perform multi-level down-sampling and semantic extraction on the features through a hierarchical encoding structure to obtain latent representations of different granularities; S3: multi-level decoding stage The latent representations are input into a multi-level Transformer decoder for hierarchical up-sampling and feature reconstruction to recover the spatial resolution of the original features layer by layer; S4: reconstruction and prediction stage The features output by the decoder are mapped back to the original feature space through a convolutional layer to obtain reconstructed features and an anomaly score map; S5: defect discrimination stage The difference between the reconstructed features and the original features is calculated, and the anomaly score map is combined to locate and discriminate defects.

2. The method of claim 1, wherein the method is based on a multi-level Transformer. Step S1 is as follows: S1.1: input the original image into the feature extraction module; Specifically, the original image is divided into defect images and non-defect images; in the training stage, only non-defect images are used to make the model learn the data distribution of non-defect samples; in the inference stage, whether to use defect images is not limited, and the model will reconstruct the input image and output the discrimination result; The training images are pre-processed: first the images are resized to pixels, then bilinear interpolation is applied to scale to pixels, padding size of pixels is applied, random rotation by degrees, random cropping to pixels, conversion to tensor format, and normalization using the mean and standard deviation of the ImageNet dataset; S1.2: use a pre-trained ResNet18 network as the backbone feature extractor; A pre-trained ResNet18 network is used as the feature extractor to extract feature maps of three different levels through a forward propagation hook function; the sizes of the three extracted feature maps are: (1) layer1 output: where B is the batch size (2) layer2 output: ; (3) layer3 output: ; S1.3: use a feature embedding concatenation function to fuse the multi-scale features; Specifically, for two feature maps with different spatial resolutions and wherein has a spatial resolution of and a number of channels of , has a spatial resolution of and a number of channels of , the process of feature embedding concatenation is as follows: (1) Calculate step size ; (2) For high-resolution feature maps Perform the unfold operation: using Parameters, will Decompose into sizes of Non-overlapping patches, output dimension is The third dimension represents the content within each patch. Each pixel is expanded into an independent channel; (3) Initialize zero tensors with dimensions ; (4) To Loop through each position and concatenate them along the channel dimension: wherein from to ; (5) the concatenated tensor is reshaped into a two-dimensional form of ; (6) Recovering spatial structure using fold operation: using parameters, reorganize features into spatial form; The mathematical representation of this process is: ; This embedding concatenation method can realize deep interaction of features of different scales while preserving multi-scale spatial information compared with simple feature concatenation or addition; S1.4: recursively apply the embedding concatenation operation to fuse the three layers of features layer by layer to obtain a fusion feature representation containing multi-scale semantic information; Specifically, the three extracted layers of features are embedded and concatenated twice: (1) Fusing features of layer 1 and layer 2 ; wherein is a layer1 feature , is a layer2 feature , the dimension of outputs_temp after fusion is ; (2) fuse the result with layer3 features again ; wherein is a layer3 feature The final fused feature outputs_final has dimension .

3. The method of claim 1, wherein the method is based on a multi-level Transformer. Step S2 is specifically as follows: first, the fusion features are mapped to the input dimension of the Transformer through convolution ; global position encoding is applied to the feature map , which adds position information to each spatial position of the feature map in a learnable embedding manner: ; wherein are learnable position embedding parameters; The feature map is flattened into a sequence form, and the dimension is transformed into wherein is a batch size, and is a height and a width of the feature map, is a number of channels; For each coding level the following is performed: (1) the input features are reorganized by block division according to the input feature size and the output feature size; (2) add a learnable zero vector as a global token before each patch; (3) Apply local position encoding Add relative position information for each position within the patch; (4) perform self-attention calculation and feedforward network processing through the Transformer Encoder; (5) separate global token and local token, global token is passed to next level as the abstract representation of this level, local token is passed through the bottleneck layer saved after compression for the decoding phase; The TransformerEncoder includes a multi-head self-attention mechanism and a feedforward neural network, and the calculation formula of the multi-head self-attention mechanism is: ; wherein, , , are a query, key, value matrix, respectively, is a transpose of the matrix, is a key vector dimension, is a number of attention heads, is an output projection matrix; The feedforward neural network uses two fully connected layers with an activation function in between: ; wherein is a GELU or ReLU activation function, , is a weight matrix, , is a bias vector; bottleneck layer The following structure is used, consisting of two linear layers and layer normalization in the middle, using a GELU activation function: ; wherein reducing the dimension from compressing to restoring the dimension to The structure is used to compress the local features of the encoder output, reducing the storage overhead while preserving the key semantic information for feature reconstruction in the decoding stage.

4. The method of claim 1, wherein the method is based on a multi-level Transformer. Step S3 is implemented as follows: starting from the deepest level, the decoding operation is performed layer by layer upwards; for each decoding level the following operations are performed: (1) concatenating the global token for this level and the local token saved at the encoding stage in the sequence dimension as memory input for the decoder; (2) using learnable query embeddings as input to the decoder The number of query embeddings is equal to ; (3) perform cross-attention calculation through the TransformerDecoder to make the query vector focus on the features in the encoding stage; (4) reorganize the decoding output into a spatial structure to recover the feature map with higher resolution; The above operations are repeated to upsample layer by layer until the resolution of the original feature map is restored, and the final decoding output feature dimension is ; The calculation process of the TransformerDecoder includes: self-attention calculation of the target sequence, enabling the query vectors to interact with each other; encoder-decoder cross-attention calculation, enabling the query vectors to focus on the output features of the encoder; and feedforward neural network processing; In the cross-attention mechanism, from the decoder input, and from the encoder output, calculated by the formula: ; wherein is a query matrix of the decoder, and are key and value matrices, respectively, output by the encoder.

5. The method of claim 1, wherein the method is based on a multi-level Transformer. Step S4 is specifically as follows: the features output by the decoder are transformed from back to the original dimension format; the first convolutional layer maps the features back to the original number of channels of the fused features to obtain reconstructed features ; performing a detach operation on the reconstructed features to block gradient propagation, then passing through a second multi-layer convolutional network processing: (1) sequentially passing through a plurality of convolution, batch normalization, and LeakyReLU activation function (2) gradually reduce the number of channels from to , , ; (3) Finally by Convolutional output single-channel anomaly score map ; Returning to the reconstructed features and anomaly score map as model output.

6. The method of claim 1, wherein the method is based on a multi-level Transformer. Step S5 is specifically as follows: calculating the reconstruction feature from the original fusion feature norm distance :​ ; The norm calculation is performed in the channel dimension, and the spatial dimension is maintained; performing a multi-scale pooling operation on the normalized distance map: using respectively , , , average pooling; up-sampling the pooling result of each scale back to the original resolution; concatenating the features of all scales; For the first dimension, the pooling operation formula is: ; wherein ; The multi-scale features after splicing are weighted and fused: ; wherein the weights ; The final anomaly score map is upsampled to the input image resolution, and thresholding processing is performed to obtain a binary mask of the defect region; The model training process includes: Only normal samples without defects are used for training; and the loss function is defined as mean square error loss: ; wherein, is a reconstruction loss, is an anomaly score prediction loss, denotes an expectation operation; The parameter is updated using Adam optimizer with learning rate set to ; during training, the backbone network parameters are fixed and not updated, only the parameters of the Transformer encoder, decoder and convolutional layers are trained; when the loss on the validation set no longer decreases, save the model checkpoint.

Citation Information

Patent Citations

  • Road semantic segmentation method based on lightweight multi-scale fusion Transform network

    CN119445107A

  • Interaction action detection method and device based on multi-level features

    CN120388227A

  • Industrial product defect detection method based on multi-granularity feature fusion

    CN120672680A

  • Industrial product surface defect detection system based on segmentation all-in-one model

    CN120707491A

  • Layered reconstruction-based general visual unsupervised defect detection method

    CN121053098A

Cited By

  • Fuzzing and pilling rating method and device based on visual continuous regression

    CN121810682A

  • A method and device for rating fuzz and pilling based on visual continuous regression

    CN121810682B

  • Image anomaly detection method, device and equipment based on feature reconstruction

    CN122223005A