Few-shot leather anomaly detection method based on domain adversarial and multi-scale fusion
Through a multi-source domain adaptive network based on domain adversarial learning, combined with multi-scale feature fusion and global attention module, the problems of cross-domain distribution differences and small sample scenarios in leather detection are solved, and high-precision anomaly detection is achieved.
Patent Information
- Application Number
- CN202510892645.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-06-30
AI Technical Summary
In industrial leather detection, high-precision detection is difficult to achieve in scenarios such as large cross-domain distribution differences, insufficient adaptability of single-source domains and small sample sizes. The existing methods rely on a large amount of labeled data and are difficult to cope with complex and diverse target domain scenarios.
A multi-source domain adaptive network based on domain adversarial learning is adopted, combined with a multi-scale feature fusion module and a global attention module, and a multi-scale feature fusion module is used to fuse features with different resolutions, and a global attention mechanism is introduced to improve the robustness of the model to complex shapes and textures, and domain adversarial learning is used to extract domain invariant features.
It realizes accurate detection of complex and diverse target domain data under the condition of few samples, improves the detection accuracy and adaptability of the model, reduces the dependence on abnormal samples acquisition, and adapts to the effective learning of multi-source domain data.
Smart Images

Figure CN120375109B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and more specifically, relates to a multi-source domain few-sample leather anomaly detection method based on domain adversarial learning with comprehensive consideration of shape and texture, which can be used to accurately detect anomalies in few-sample leather with shape-texture anomalies in industrial visual inspection. Background Art
[0002] In industrial production, in order to ensure that products meet standards and have no production defects, it is usually necessary to inspect products on the production line. However, since the number of products is usually very large, relying solely on manual inspection is impractical.
[0003] With the development of modern industry, anomaly detection has become an integral part of the production process, and the development of machine vision technology has provided a new solution for anomaly detection. This involves installing cameras and light sources on the production line and using machine vision technology to analyze images of products on the line, automatically detecting surface anomalies and improving production efficiency and product quality.
[0004] The leather industry is a key sector within my country's light industry. Leather products are high-end consumer goods, and consumers often choose them based on their appearance and quality. However, surface defects not only affect the aesthetics of the product but also its quality and lifespan. Traditional rule-based methods, however, often struggle to meet practical needs due to their limitations when dealing with complex and diverse industrial scenarios. For example, in the inspection of materials such as leather and fabric, anomalies can manifest in a variety of shapes and textures due to the diversity and uncertainty between samples. Therefore, effectively detecting and segmenting these anomalies has become a major challenge in intelligent industrial inspection.
[0005] In recent years, the rapid development of deep learning technology, particularly convolutional neural networks (CNNs), has brought significant progress to the field of anomaly detection. Supervised deep learning methods, in particular, have achieved high-precision anomaly detection by leveraging large-scale labeled datasets. However, in real-world industrial applications, obtaining a large number of labeled anomaly samples is extremely costly, and in some scenarios, it is almost impossible to obtain sufficient anomaly samples. This limits the application of supervised learning methods. Few-shot learning, as an emerging learning paradigm, addresses this problem by learning features from a limited number of samples and has become a hot topic in the field of anomaly detection.
[0006] Another key challenge facing anomaly detection is the distribution differences in cross-domain data. In real-world scenarios, products from different sources or batches differ significantly in terms of materials, textures, and other aspects. This results in a significant performance degradation when models trained in the source domain are directly applied to the target domain. To address this issue, domain adaptation (DA) methods have been proposed, aiming to improve the generalization ability of the model in the target domain by reducing the distribution differences between the source and target domains. In particular, the Multi-Source Domain Adaptation (MSDA) method, by combining knowledge from multiple source domains, can significantly improve the model's adaptability to complex and diverse target domain data, demonstrating better robustness and generalization capabilities than traditional single-source domain adaptation methods.
[0007] To sum up, the existing technologies have the following problems and defects: First, there are significant cross-domain distribution differences. Products from different sources or batches have large differences in material, texture and other features, which leads to a serious performance degradation of the source domain model when directly applied to the target domain; second, the single source domain adaptability is insufficient. Traditional domain adaptation methods rely on a single source domain and have difficulty coping with complex and diverse target domain scenarios, and their generalization capabilities are limited; finally, there are challenges in few-sample scenarios. Abnormal samples are scarce in industrial scenarios. Existing methods rely on a large amount of labeled data, making it difficult to achieve high-precision detection under few-sample conditions. Summary of the Invention
[0008] To overcome the shortcomings of the aforementioned prior art, this paper proposes a multi-source domain few-shot leather anomaly detection method based on domain adversarial learning, integrating shape and texture information. This method includes a novel multi-source domain adaptive network that comprehensively considers both texture and shape information, aiming to address the problem of few-shot anomaly detection in industrial scenarios. Based on domain adversarial learning, this method designs three proxy tasks to learn shape features, texture features, and a combination of these features, and then achieves joint optimization. By introducing an adversarial mechanism, domain adversarial learning enables the model to not only extract domain-invariant features but also better adapt to anomaly detection tasks in the target domain.
[0009] To achieve effective feature fusion and improve the model's detection accuracy, this paper introduces two key modules: a multi-scale feature fusion module and a global attention module. The design of the multi-scale feature fusion module is inspired by multi-scale feature extraction methods such as pyramid networks. By fusing features from different resolutions, this module enables the model to simultaneously focus on global responses to large-scale anomalies and detailed detection of small-scale defects. The global attention module, by introducing a global attention mechanism, increases the focus on important areas during the decoding phase, making the model more robust and accurate when handling complex shapes and textures.
[0010] In practical applications of anomaly detection, the diversity and authenticity of the dataset have a crucial impact on the performance of the model. To better evaluate the proposed multi-source domain adaptation network, we constructed a leather anomaly detection dataset containing rich texture and shape information and conducted systematic experimental validation on the dataset.
[0011] The technical solution of the present invention includes the following: a small sample leather anomaly detection method based on domain adversarial and multi-scale fusion, comprising the following steps:
[0012] Step 1: Construct abnormal image training dataset and leather dataset;
[0013] Step 2: Build a dual encoder-decoder network based on multi-scale and shape-texture information fusion; each encoder of the dual encoder includes multiple Transformers for extracting features of different scales, and a multi-scale feature fusion module for fusing the features of different scales extracted by different encoders. The decoder includes multiple global attention modules, where the multi-scale feature fusion module includes a convolutional layer, an SE module, a spatial attention module, a channel attention mechanism, and a local attention mechanism. The global attention module includes a channel attention mechanism and a convolutional layer.
[0014] Step 3: Build a domain adversarial learning network, which adds a multi-class domain discriminator to the dual encoder-decoder network based on multi-scale and shape-texture information fusion;
[0015] Three different proxy tasks are constructed for shape features, texture features, and shape-texture features. Different multi-source data are randomly selected from the abnormal image training dataset and the leather dataset to perform alternating training on the three proxy tasks.
[0016] Step 4: Use the trained dual encoder-decoder network based on multi-scale and shape-texture information fusion to achieve accurate anomaly detection on the leather dataset to be tested.
[0017] Furthermore, the specific processing process of the dual encoder-decoder network based on multi-scale and shape-texture information fusion in step 2 is as follows:
[0018] The input image is first adjusted to a high-resolution image Rgb and a low-resolution image Srgb. These two are used as the two inputs of the dual encoder. The high-resolution image Rgb and the low-resolution image Srgb are processed in multiple stages by Transformer. Feature extraction is performed at each stage to obtain feature and , where i represents the i-th stage and also the i-th scale; at the same time, on this basis, the features of each stage are fused by the multi-scale feature fusion module and Perform feature fusion and transmit the fused features to the decoder.
[0019] Furthermore, the processing process of the multi-scale feature fusion module is as follows: Perform upsampling to restore to the same Same size, then use convolutional layers to reduce and Dimensions:
[0020]
[0021] in represents bicubic interpolation upsampling, represents a convolutional layer with a kernel size of 1×1, and It is the multi-scale feature after dimensionality reduction, and then weighted by SE module and The features of each channel are then extracted through the spatial attention module to extract the spatial attention map and :
[0022]
[0023] Among them, SE represents the channel feature weighting operation, and SA represents the spatial attention operation;
[0024] In order to guide the fusion of the features of the low-resolution Srgb image, the spatial attention map of the high-resolution Rgb and Perform element-wise multiplication and then add Similarly, in order to guide the fusion of high-resolution RGB image features, the low-resolution SRGB spatial attention map is added. and Perform element-wise multiplication and then add Finally, the two are restored to dimension through the convolution layer and the multi-scale features are fused by element addition to obtain :
[0025]
[0026] exist Based on Each channel of the feature map is weighted, and then the local attention mechanism is used to locally weight the weighted feature map, and finally the result of multi-scale feature fusion F is obtained. i :
[0027]
[0028] Among them, LA represents the local attention mechanism, CA represents the channel attention mechanism, Represents the final multi-scale fused feature map.
[0029] Furthermore, the specific processing of the global attention module is as follows;
[0030] The current layer features fused by the multi-scale feature fusion module and the output of the previous global attention module in the decoder As input, the features Perform channel attention weighting, since and The corresponding number of channels is different, for features Perform upsampling to restore the number of channels to the same level as the feature Same, finally add the two elements and perform convolution operation to get :
[0031]
[0032] in, represents the feature output of the decoder, represents a 1×1 convolution, represents bicubic interpolation upsampling, represents the channel attention weight.
[0033] Furthermore, in step (3), the dual encoder is used as a feature generator to extract features from the input image and obtain multi-scale fusion features, which are then fed into a multi-class domain discriminator;
[0034] The multi-class domain discriminator is connected to the dual encoder using a gradient reversal layer. Specifically, during forward propagation, the gradient reversal layer does not change the transmission of features, but during backpropagation, the gradient sign of the multi-class domain discriminator is reversed before being passed to the dual encoder.
[0035] Furthermore, the first proxy task contains three data sets, which are respectively 、 and In this dataset, and are two source domain datasets, is the target domain dataset. In each training cycle, two original abnormal image training datasets are randomly selected as the source domain and , and the remaining leather dataset is set as the target domain ;
[0036] The total loss for the first proxy task is:
[0037]
[0038]
[0039]
[0040] in, represents the total loss of the first proxy task, represents the loss of source domain data training, represents the segmentation loss of the target domain data, Indicates that the source domain data The segmentation loss, Represents the source domain data in the first agent task The confrontation loss, It is a hyperparameter that adjusts the relative importance of segmentation loss and adversarial loss; represents the segmentation output of the decoder layer i, , is the convolution operation, is the activation function, represents the binary cross entropy loss, Represents the true value, that is, the actual segmentation label.
[0041] Furthermore, the second proxy task also contains three data sets, which are respectively 、 and In this dataset, and are two source domain datasets, is the target domain dataset. In each training cycle, two original abnormal image training datasets are randomly selected as the source domain and , and the remaining leather dataset is set as the target domain ;
[0042] The total loss for the second proxy task is:
[0043]
[0044]
[0045]
[0046] in, represents the total loss of the second proxy task, represents the loss of source domain data training, represents the segmentation loss of the target domain data, Indicates that the source domain data The segmentation loss, Represents the source domain data in the second agent task The confrontation loss, It is a hyperparameter that adjusts the relative importance of segmentation loss and adversarial loss; represents the segmentation output of the decoder layer i, , is the convolution operation, is the activation function, represents the binary cross entropy loss, Represents the true value, that is, the actual segmentation label.
[0047] Furthermore, the third agent task takes all the data of the first and second agent tasks as input;
[0048] The total loss for the third proxy task is:
[0049]
[0050] in, represents the total loss of the third proxy task, Indicates that the source domain data and segmentation loss of target domain data.
[0051] Furthermore, the total loss of domain adversarial learning network training is obtained by weighting the losses of the three proxy tasks. During training, the gradient of the loss function with respect to the model parameters is calculated through the backpropagation algorithm, and the model parameters are updated using the optimization algorithm.
[0052] Furthermore, the model parameter optimization is divided into two stages: for the first proxy task, the first stage comes from The images are used to optimize the model:
[0053]
[0054] in is the learning rate, is the parameter optimized by random alternation training of the first agent task, Represents the gradient of the loss function with respect to the parameters, which is used to indicate the optimization direction. The first proxy task targets the source domain dataset The loss function of training; in the second stage, similarly, using the target domain dataset The image in the target domain is segmented by the loss Update model parameters ; are the updated model parameters;
[0055] The same model parameter optimization method is used to implement the second and third proxy tasks.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] The present invention is not subject to the high cost of obtaining abnormal samples. It can achieve accurate detection by combining a large amount of data rich in shape and texture information with a small number of labeled target domain abnormal samples. The multi-source domain adaptive network uses three proxy tasks to achieve effective learning and optimization of multi-source domain data, significantly improving the model's adaptability to complex and diverse target domain data. At the same time, through the introduction of a multi-scale feature fusion module and a global attention module, effective feature fusion is achieved and the detection accuracy of the model is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is the overall technical flow chart of the present invention.
[0059] Figure 2 Schematic diagram of random alternating training groups in an embodiment of the present invention.
[0060] Figure 3 FIG. 4 is an overall structural diagram of a multi-source domain adaptive network in an embodiment of the present invention.
[0061] Figure 4 This is a structural diagram of the dual encoder-decoder network based on multi-scale and shape-texture information fusion, the multi-scale feature fusion module, and the global attention module in an embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to better understand the technical solution of the present invention, the specific embodiments of the present invention are further described below with reference to the accompanying drawings.
[0063] Reference Figure 1 As shown, the embodiment of the present invention provides a multi-source domain few-sample leather anomaly detection method based on domain adversarial learning with comprehensive consideration of shape and texture. The overall implementation process is as follows:
[0064] (1) Construct an abnormal image training dataset and a leather dataset. The abnormal image training dataset uses the publicly available MVTec AD dataset for anomaly detection. The MVTec AD dataset is a dataset widely used in industrial visual inspection tasks, containing various categories of industrial products and their corresponding normal and abnormal samples. The data for each category includes an RGB image and its corresponding annotation information (such as a mask). The leather dataset is a self-constructed dataset with a small amount of annotations. The abnormal samples include defects such as scratches and texture anomalies.
[0065] The three datasets involved in the network's first proxy task, shape feature learning, are the first random alternating training set. This includes two images with rich shape information and significant shape differences randomly selected from the public MVTec AD dataset, and a leather image with a small amount of annotations randomly selected from a self-built leather dataset. These serve as the two source domain datasets and target domain datasets for the network's first proxy task, shape feature learning.
[0066] The second proxy task in the network, texture feature learning, is similar to the previous one. The difference is that the source domain dataset selected in the second random alternating training group is two randomly selected images from the public dataset MVTec AD that are rich in texture information and have significant shape differences.
[0067] The third proxy task in the network, the integrated learning of shape and texture features, takes all the data from proxy tasks one and two as input, allowing the model to integrate and learn the overall data features again.
[0068] (2) Build a dual encoder-decoder network based on multi-scale and shape-texture information fusion. The dual encoder-decoder network consists of multiple Transformers, a multi-scale feature fusion module (equivalent to a dual encoder), and a global attention module. The multi-scale feature fusion module includes a convolutional layer, a Squeeze-and-Excitation (SE) module, a spatial attention (CBAM: Convolutional Block Attention Module) module, a channel attention mechanism, and a local attention mechanism. The global attention module includes a channel attention mechanism and a convolutional layer.
[0069] The specific implementation of step (2) is as follows:
[0070] Reference Figure 3 and Figure 4As shown in the figure, the original image, i.e., the image selected from the abnormal image training dataset and the self-built leather dataset, will be adjusted to a high-resolution image RGB and a low-resolution image SRGB. These two serve as the two inputs of the dual encoder. Transformer adopts the MiT (Mix Vision Transformer) architecture that combines convolution and self-attention mechanisms. Its main advantage is that it can efficiently extract local and global feature information. The convolution layer is used to retain the local features of the input image, especially low-level features such as texture and edges. The multi-head self-attention mechanism (MSA) captures long-range dependencies through global modeling and extracts global information. MiT uses block attention and linear complexity attention mechanisms to reduce computational overhead. Transformer processes the input image in multiple stages, and each stage extracts features from it. and Indicates that, on this basis, the features of each stage are fused by the multi-scale feature fusion module and Perform feature fusion and transmit the fused features to the decoder;
[0071] The principle of the multi-scale feature fusion module is to use the attention map of the low-resolution SRGB image to guide the global response of large-scale anomalies in the high-resolution RGB image. Similarly, the attention map of the high-resolution RGB image guides more detailed detection of small-scale defects in the low-resolution SRGB image.
[0072] The features extracted in the feature extraction stage of each layer of the encoder and It is used as the input of the multi-scale feature fusion module. Since Srgb is compressed from Rgb, the features extracted are The resolution of is low, so it is necessary to Perform upsampling to restore to the same For the same size, in the subsequent feature fusion, the convolution layer is first used to reduce and dimensionality to reduce computational overhead and enhance the expression of important features:
[0073]
[0074] in Represents bicubic interpolation upsampling. Compared with bilinear interpolation, bicubic interpolation can better preserve feature details by calculating smoother pixel value transitions. represents a convolutional layer with a kernel size of 1×1, and It is the multi-scale feature after dimensionality reduction, and then the SE (Squeeze-and-Excitation) module weights and The features of each channel are then extracted through the spatial attention module (CBAM: Convolutional Block AttentionModule) and :
[0075]
[0076] Among them, SE represents the channel feature weighting operation, and SA represents the spatial attention operation;
[0077] The SE (Squeeze-and-Excitation) module first performs channel compression, that is, generates global information in the channel dimension through a global average pooling operation, and then generates weights for each channel through channel activation, that is, through two fully connected layers and nonlinear activation (such as ReLU and Sigmoid). Finally, the generated weights are used to weight the features of each channel, strengthening important features and suppressing irrelevant or redundant information.
[0078] The spatial attention module first performs spatial attention calculations, generates attention weights in the spatial dimension by performing global maximum pooling and average pooling on the feature map, and then weights the feature map, that is, uses the attention map to weight the original feature map to highlight the salient areas.
[0079] In order to guide the fusion of the features of the low-resolution Srgb image, the spatial attention map of the high-resolution Rgb and Perform element-wise multiplication and then add Similarly, in order to guide the fusion of high-resolution RGB image features, the low-resolution SRGB spatial attention map is added. and Perform element-wise multiplication and then add Finally, the two are restored to dimension through the convolution layer and the multi-scale features are fused by element addition to obtain :
[0080]
[0081] exist Based on the channel attention mechanism (CA), Each channel of the feature map is weighted, and then the local attention mechanism (LA) is used to locally weight the weighted feature map, and finally the multi-scale feature fusion result F is obtained. i :
[0082]
[0083] Among them, LA represents the local attention mechanism, CA represents the channel attention mechanism, Represents the final multi-scale fused feature map;
[0084] The main goal of the channel attention mechanism (CA) is to enhance the expressiveness of important feature channels while suppressing redundant information by weighting the channel dimensions of the feature map. The specific process is as follows:
[0085] First, the input feature map F (with a size of H×W×C) is globally average pooled in the spatial dimension to obtain a 1×1×C global description vector. This vector represents the global semantic information of each channel. Subsequently, the vector passes through a multi-layer perceptron (MLP) consisting of two fully connected layers. In the first fully connected layer, the number of channels is reduced to 1 / 4 of the original number of channels to reduce computational complexity. Next, nonlinear mapping is performed using the ReLU activation function, and then the original number of channels is restored through the second fully connected layer. Finally, the weight of each channel is generated using the Sigmoid activation function. Finally, the generated weight is multiplied element-by-element by each channel of the input feature map to complete the weighting of the feature channels. This process can be expressed as:
[0086] ,
[0087] Where GAP represents global average pooling, Represents the Sigmoid function, ⊙ represents channel-by-channel multiplication, is the weighted feature map.
[0088] The purpose of the local attention mechanism LA is to enhance the local area of the feature map in the spatial dimension, so that the network can pay more attention to fine-grained features and small-scale abnormal areas. The specific process is as follows:
[0089] First, the input feature map F is divided into several small local regions (e.g., 7×7 windows) through a local window partitioning operation. For each local region, global maximum pooling and global average pooling are calculated to extract local saliency information. These pooling results are concatenated along the channel dimension and passed through a small convolutional network (3×3 convolutional layer) to generate a local attention weight map. (Size is H×W×1). Subsequently, this attention weight map is element-wise multiplied with the original feature map F to complete the weighting of the local area. This process can be expressed as:
[0090] ,
[0091] in is the local attention weight, is the weighted feature map.
[0092] Through the local attention mechanism, the network can capture detailed areas more accurately, which helps to detect small-scale anomalies or subtle features.
[0093] Since the features extracted by each layer of encoder are fused by the multi-scale feature fusion module, the feature results The number of channels is different. For this reason, a new module global attention module is proposed and integrated into the decoder. The specific processing process of the global attention module is as follows;
[0094] The current layer features fused by the multi-scale feature fusion module and the output of the previous global attention module in the decoder As input, the features Perform channel attention weighting, since and The corresponding number of channels is different, so we need to Perform bicubic interpolation upsampling to adjust the number of channels to match the feature Match, and finally add the two element by element and perform convolution operation to get :
[0095]
[0096] in, represents the feature output of the decoder, represents a 1×1 convolution, represents bicubic interpolation upsampling, represents the channel attention weight;
[0097] (3) Building a domain adversarial learning network, wherein the domain adversarial learning network adds a multi-class domain discriminator to the dual encoder-decoder network based on multi-scale and shape-texture information fusion described in step (2);
[0098] The specific implementation of step (3) is as follows:
[0099] Reference Figure 3As shown in the figure, the dual encoder is first used as a feature generator to extract features from the input image. The input data includes two source domain datasets: two randomly selected datasets with rich shape information and significant shape differences from the public MVTec AD dataset, and two randomly selected datasets with rich texture information and significant shape differences; and a target domain dataset: a randomly selected leather dataset with a small amount of annotations. The dual encoder generates multi-scale fused features for these three types of data through its feature extraction module, the Transformer, and its multi-scale feature fusion module. These features contain information from both high-resolution RGB images and low-resolution RGB images, capturing both the shape and texture characteristics of the target.
[0100] The fused features of the two source domain datasets are then fed into a multi-class domain discriminator. The multi-class domain discriminator is a classification network whose goal is to classify input features to determine their domain of origin. Specifically, the multi-class domain discriminator outputs a probability distribution for each feature belonging to each source domain and calculates the classification error using a cross-entropy loss function to ensure that the discriminator can accurately identify the feature's domain of origin.
[0101] To achieve domain adversarial learning, a gradient reversal layer is used to connect the multi-class domain discriminator to the dual encoder. The role of the gradient reversal layer is to reverse the sign of the gradient during backpropagation, so that the discriminator and the dual encoder form an adversarial relationship. Specifically, during forward propagation, the gradient reversal layer does not change the transmission of features, but during backpropagation, the sign of the gradient of the multi-class domain discriminator is reversed before being passed to the dual encoder. Through this adversarial mechanism, when extracting features, the dual encoder not only focuses on the task requirements of segmentation defects, but also confuses the judgment of the domain discriminator to the greatest extent, thereby forcing the dual encoder to extract domain-invariant features that are independent of the source domain and only relevant to the target task. These domain-invariant features help improve the model's generalization ability on few-sample target domain data, resulting in stronger segmentation performance in the target domain.
[0102] The overall framework of the domain adversarial learning network realizes the combination of feature extraction and domain adversarial learning, which not only retains the effectiveness of the dual encoder-decoder network in anomaly detection, but also enhances the robustness and domain invariance of features through multi-class domain discriminators and gradient reversal layers.
[0103] (4) The steps of multi-source domain random alternating training for shape feature learning, texture feature learning, and shape-texture feature integrated learning are similar to those for shape feature learning. The difference lies in the different training data and the fact that domain adversarial learning is not performed in shape-texture feature integrated learning. Specifically, the steps include the following:
[0104] (4a) Preprocess the shape anomaly images in the original anomaly dataset and the leather anomaly images in the target domain leather dataset. The specific implementation is as follows:
[0105] Reference Figure 2 As shown in Figure 2, since we use multi-source domain random alternation training, at the beginning of each training cycle, the source domain and target domain settings of each proxy task are randomized to enhance the generalization ability of the model. For the first proxy task shape feature learning, a random alternation training group is used, which contains three data sets, respectively denoted as 、 and In this dataset, and are two source domain datasets, is the target domain dataset. In each training cycle, two original abnormal datasets are randomly selected as source domains and , and the remaining leather dataset is set as the target domain This random selection is repeated in each training cycle to ensure that the model is exposed to different combinations of source and target domains during training, improving its ability to learn domain-invariant features.
[0106] Next, to meet the requirements of the shape anomaly detection task, we preprocessed the shape anomaly images from the original anomaly dataset and the leather anomaly images from the target domain leather dataset. Specifically, these images were resized to two resolutions: a high-resolution image (Rgb) and a low-resolution image (Srgb). The resolution of the high-resolution image (Rgb) was set to 384×384 to fully preserve the global information of the anomaly. The low-resolution image (Srgb) was obtained by downsampling the high-resolution image (Rgb) by a factor of two, to a resolution of 192×192. This high- and low-resolution image construction was designed to fully utilize the dual encoder's ability to combine global and local feature information.
[0107] The processed high-resolution RGB image and low-resolution SRGB image serve as the two inputs of the dual encoder, extracting shape features at different resolutions. This paves the way for subsequent feature fusion and shape feature learning for proxy tasks. This preprocessing approach ensures diverse input data and fully utilizes multi-scale features, providing the model with greater adaptability in shape anomaly detection tasks.
[0108] (4b) The encoder in the dual encoder-decoder network based on multi-scale and shape-texture information fusion extracts features from the shape abnormality image and the leather abnormality image and performs feature fusion, and transmits the final extracted fusion features to the decoder. The specific implementation method is as follows:
[0109] Reference Figure 3 and Figure 4 As shown, first, the two inputs Rgb and Srgb of the dual encoder are subjected to four-stage feature extraction through the Transformer module to obtain and , and then the features extracted in each stage and The multi-scale fusion feature map is obtained by feature fusion through the multi-scale feature fusion module , the feature map after the fusion of these four stages It integrates multi-scale and shape information and can more comprehensively express abnormal features in images.
[0110] (4c) The multi-class domain discriminator in the domain adversarial learning network is used to distinguish which source domain data the features generated by the dual encoder come from. The gradient reversal layer is used to implement adversarial training, so that the dual encoder, as a feature extractor, can "cheat" the discriminator during the optimization process, thereby enabling the dual encoder to extract domain-invariant features. The specific implementation method is as follows:
[0111] Reference Figure 3 As shown, first the two source domains extracted by the dual encoder are and The features of are used as input and passed to the multi-class domain discriminator for classification. The multi-class domain discriminator is essentially a binary classifier, consisting of a small fully connected network with a simple and efficient structure. The output layer of the discriminator consists of a single node, and the output value is whether the input feature belongs to the source domain. Through this probability value, the discriminator tries to distinguish the features generated by the dual encoder as belonging to the source domain according to their source. or source domain For example, if the output value is close to 1, the discriminator believes that the feature comes from the source domain ; If the output value is close to 0, the feature is considered to be from the source domain .
[0112] During the training process, in order to make the features generated by the dual encoder difficult for the discriminator to correctly classify, a gradient reversal layer is used to implement adversarial learning. The gradient reversal layer flips the sign of the gradient returned from the discriminator during back propagation, thereby making the dual encoder "cheat" the discriminator during the optimization process, forcing it to extract features that are as difficult to distinguish as those from the discriminator. or , thereby realizing the extraction of domain-invariant features.
[0113] On this basis, the adversarial loss function in adversarial learning is defined , used to optimize multi-class domain adversarial learning. Specifically, the loss function uses the cross entropy loss function form, expressed as:
[0114]
[0115] Where K represents the number of source domains. In this task, K=2. Represents the domain label of the source domain sample k. If the sample comes from the source domain ,but =1; if the sample comes from the source domain ,but =0. Indicates that the multi-class domain discriminator predicts that the source domain sample k belongs to the source domain By minimizing this adversarial loss function, the classification ability of the multi-class domain discriminator can be optimized. At the same time, the dual encoder is reversely optimized through the gradient reversal layer, making the features extracted by the dual encoder difficult to distinguish by the discriminator, thus achieving domain invariance.
[0116] Through this process, the dual encoder not only extracts high-quality features for defect segmentation, but also extracts domain-specific common features, enhancing its generalization across multi-source domain data. Ultimately, this adversarial learning mechanism effectively improves the dual encoder's ability to extract domain-invariant features, providing more robust feature support for subsequent tasks.
[0117] (4d) The global attention module in the decoder of the dual encoder-decoder network based on multi-scale and shape-texture information fusion is used to fully integrate the features extracted by the encoder at each layer and make output predictions. The final prediction result is compared with the ground truth (GT) for loss calculation. The gradient of the loss function with respect to the model parameters is calculated using the backpropagation algorithm, and the model parameters are updated using the optimization algorithm. The above steps are repeated until all training rounds are completed. The specific implementation method is as follows:
[0118] Reference Figure 3 As shown in the figure, the features of each stage extracted by the dual encoder are input into the global attention module of the decoder to obtain the feature output of each decoder stage. In order to improve the accuracy of training detection, we perform output prediction at each decoder stage: Subsequently, in order to predict the segmentation results at each stage, a layer-by-layer output prediction mechanism is designed. For the decoder output features of each stage , through the convolution operation and activation function Get the segmentation output of this stage , the calculation formula is:
[0119]
[0120] in, represents the segmentation output of the decoder layer i, Indicates the method used to obtain the final segmentation result Function, used to map the output result to the interval [0,1] as the probability value of the final segmentation result, and finally calculate the loss between the predicted result and the corresponding segmentation label ground truth (GT):
[0121]
[0122]
[0123] in, is the outlier segmentation loss, Indicates the decoder's outputs, represents the binary cross-entropy (BCE) loss, is the total number of pixels, represents the index of the pixel, and That is the The ground truth and segmentation results of pixels. Since the segmentation output of the last layer integrates all the previous features, we use it as the final segmentation result in the test phase;
[0124] Therefore, the total loss of the first proxy task is:
[0125]
[0126]
[0127]
[0128] in, represents the total loss of the first proxy task, represents the loss of source domain data training, represents the segmentation loss of the target domain data, Indicates that the source domain data The segmentation loss, Represents the source domain data in the first agent task The confrontation loss, It is a hyperparameter that adjusts the relative importance of segmentation loss and adversarial loss;
[0129] Similarly, the total loss of the second proxy task is:
[0130]
[0131]
[0132]
[0133] in, represents the total loss of the second proxy task, represents the loss of source domain data training, represents the segmentation loss of the target domain data, Indicates that the source domain data The segmentation loss, Represents the source domain data in the second agent task The confrontation loss, It is a hyperparameter that adjusts the relative importance of segmentation loss and adversarial loss;
[0134] Similarly, the total loss of the third proxy task is similar, but the third proxy task has no adversarial loss, so it is:
[0135]
[0136] in, represents the total loss of the third proxy task, Indicates source domain data and target domain data Segmentation loss;
[0137] The total loss of the final network task is the sum of the losses of the three proxy tasks:
[0138]
[0139] in, 、 and is a hyperparameter that adjusts the importance of each agent task relative to other agent tasks;
[0140] During training, the gradient of the loss function with respect to the model parameters is calculated using the backpropagation algorithm, and the model parameters are updated using the optimization algorithm. The training process repeats the above steps of feature extraction, segmentation loss calculation, and parameter update until all training rounds are completed.
[0141] The model parameter optimization is divided into two stages. For the first agent task, the first stage comes from The images are used to optimize the model:
[0142]
[0143] in is the learning rate, is the parameter optimized by random alternation training of the first agent task, Represents the gradient of the loss function with respect to the parameters, which is used to indicate the optimization direction. The first agent task is to The loss function of training. In the second stage, similarly, using the data set The image in, through the loss function Update model parameters , so that it is in the dataset It also has good robustness. is the updated model parameter. In this way, we enable the model to have a good anomaly segmentation effect not only on the source domain, but also on the target domain.
[0144] The same model parameter optimization method is used to implement the second and third proxy tasks.
[0145] Finally, under the random alternating training of multiple source domains, the model can not only accurately segment anomalies but also achieve robust generalization of cross-domain data.
[0146] Step (7) uses the dual encoder-decoder network based on multi-scale and shape-texture information fusion described in step (2) to achieve accurate anomaly detection on the self-built leather dataset. The specific implementation method is as follows:
[0147] Reference Figure 3 As shown in the figure, the self-built leather dataset is input as a test set into the trained network framework, and finally the mask map for anomaly detection is obtained. The mask map for anomaly detection is processed by the sigmoid function to obtain a normalized probability map. Pixels with probability values close to 1 are considered to be abnormal areas, while pixels close to 0 represent normal areas. In this way, the trained network can accurately detect abnormal areas in the self-built leather dataset and output clear anomaly mask maps, providing reliable support for subsequent industrial quality inspection or application scenarios. The experimental data of the present invention are shown in Table 1 below. In the table, I-AUROC is mainly used to evaluate the anomaly detection performance at the image level; P-AUROC is mainly used to evaluate the anomaly localization performance at the pixel level. It can be seen from the table that the present invention has achieved outstanding results in both anomaly detection and anomaly localization.
[0148] Table 1 Detection (I-AUROC) and localization (P-AUROC)
[0149]
[0150] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Persons skilled in the art may make various modifications, additions, or substitutions to the described specific embodiments without departing from the spirit of the present invention or exceeding the scope of the appended claims.
Claims
1. A few-sample leather anomaly detection method based on domain adversarial and multi-scale fusion, which is characterized by: The steps include: Step 1: Construct abnormal image training dataset and leather dataset; Step 2: Build a dual encoder-decoder network based on multi-scale and shape-texture information fusion; each encoder of the dual encoder includes multiple Transformers for extracting features of different scales, and a multi-scale feature fusion module for fusing the features of different scales extracted by different encoders. The decoder includes multiple global attention modules, where the multi-scale feature fusion module includes a convolutional layer, an SE module, a spatial attention module, a channel attention mechanism, and a local attention mechanism. The global attention module includes a channel attention mechanism and a convolutional layer. Step 3: Build a domain adversarial learning network, which adds a multi-class domain discriminator to the dual encoder-decoder network based on multi-scale and shape-texture information fusion; Three different proxy tasks are constructed for shape features, texture features, and shape-texture features. Different multi-source data are randomly selected from the abnormal image training dataset and the leather dataset to perform alternating training on the three proxy tasks. Step 4: Use the trained dual encoder-decoder network based on multi-scale and shape-texture information fusion to achieve accurate anomaly detection on the leather dataset to be tested.
2. The method for detecting leather anomalies with a small number of samples based on domain adversarial and multi-scale fusion according to claim 1 is characterized in that: The specific processing process of the dual encoder-decoder network based on multi-scale and shape-texture information fusion in step 2 is as follows: The input image is first adjusted to a high-resolution image Rgb and a low-resolution image Srgb. These two are used as the two inputs of the dual encoder. The high-resolution image Rgb and the low-resolution image Srgb are processed in multiple stages by Transformer. Feature extraction is performed at each stage to obtain feature and , where i represents the i-th stage and also the i-th scale; at the same time, on this basis, the features of each stage are fused by the multi-scale feature fusion module and Perform feature fusion and transmit the fused features to the decoder.
3. The method for detecting leather anomalies with a small number of samples based on domain adversarial and multi-scale fusion according to claim 2, wherein: The processing process of the multi-scale feature fusion module is as follows: Perform upsampling to restore to the same Same size, then use convolutional layers to reduce and Dimensions: ; in represents bicubic interpolation upsampling, represents a convolutional layer with a kernel size of 1×1, and It is the multi-scale feature after dimensionality reduction, and then weighted by SE module and The features of each channel are then extracted through the spatial attention module to extract the spatial attention map and : ; Among them, SE represents the channel feature weighting operation, and SA represents the spatial attention operation; In order to guide the fusion of the features of the low-resolution Srgb image, the spatial attention map of the high-resolution Rgb and Perform element-wise multiplication and then add Similarly, in order to guide the fusion of high-resolution RGB image features, the low-resolution SRGB spatial attention map is added. and Perform element-wise multiplication and then add Finally, the two are restored to dimension through the convolution layer and the multi-scale features are fused by element addition to obtain : ; exist Based on Each channel of the feature map is weighted, and then the local attention mechanism is used to locally weight the weighted feature map, and finally the result of multi-scale feature fusion F is obtained. i : ; Among them, LA represents the local attention mechanism, CA represents the channel attention mechanism, Represents the final multi-scale fused feature map.
4. The method for detecting leather anomalies with a small number of samples based on domain adversarial and multi-scale fusion as claimed in claim 3 is characterized in that: The specific processing process of the attention module is as follows; The current layer features fused by the multi-scale feature fusion module and the output of the previous global attention module in the decoder As input, the features Perform channel attention weighting, since and The corresponding number of channels is different, for features Perform upsampling to restore the number of channels to the same level as the feature Same, finally add the two elements and perform convolution operation to get : ; in, represents the feature output of the decoder, represents a 1×1 convolution, represents bicubic interpolation upsampling, represents the channel attention weight.
5. The method for detecting leather anomalies with a small number of samples based on domain adversarial and multi-scale fusion according to claim 1, wherein: In step (3), the dual encoder is used as a feature generator to extract features from the input image and obtain multi-scale fusion features, which are then fed into a multi-class domain discriminator; The multi-class domain discriminator is connected to the dual encoder using a gradient reversal layer. Specifically, during forward propagation, the gradient reversal layer does not change the transmission of features, but during backpropagation, the gradient sign of the multi-class domain discriminator is reversed before being passed to the dual encoder.
6. The method for detecting leather anomalies with a small number of samples based on domain adversarial and multi-scale fusion according to claim 1, wherein: The first proxy task contains three datasets, denoted as 、 and In this dataset, and are two source domain datasets, is the target domain dataset. In each training cycle, two original abnormal image training datasets are randomly selected as the source domain and , and the remaining leather dataset is set as the target domain ; The total loss for the first proxy task is: ; ; ; in, represents the total loss of the first proxy task, represents the loss of source domain data training, represents the segmentation loss of the target domain data, Indicates that the source domain data The segmentation loss, Represents the source domain data in the first agent task The confrontation loss, It is a hyperparameter that adjusts the relative importance of segmentation loss and adversarial loss; represents the segmentation output of the decoder layer i, , is the convolution operation, is the activation function, represents the binary cross entropy loss, Represents the true value, that is, the actual segmentation label.
7. The method for detecting leather anomalies with a small number of samples based on domain adversarial and multi-scale fusion according to claim 1, wherein: The second proxy task also contains three data sets, which are respectively 、 and In this dataset, and are two source domain datasets, is the target domain dataset. In each training cycle, two original abnormal image training datasets are randomly selected as the source domain and , and the remaining leather dataset is set as the target domain ; The total loss for the second proxy task is: ; ; ; in, represents the total loss of the second proxy task, represents the loss of source domain data training, represents the segmentation loss of the target domain data, Indicates that the source domain data The segmentation loss, Represents the source domain data in the second agent task The confrontation loss, It is a hyperparameter that adjusts the relative importance of segmentation loss and adversarial loss; represents the segmentation output of the decoder layer i, , is the convolution operation, is the activation function, represents the binary cross entropy loss, Represents the true value, that is, the actual segmentation label.
8. The method for detecting leather anomalies with a small number of samples based on domain adversarial and multi-scale fusion according to claim 1, wherein: The third agent task takes all the data of the first and second agent tasks as input; The total loss for the third proxy task is: ; in, represents the total loss of the third proxy task, Indicates that the source domain data and segmentation loss of target domain data.
9. The method for detecting leather anomalies with a small number of samples based on domain adversarial and multi-scale fusion as claimed in claim 1, characterized in that: The total loss of adversarial learning network training is obtained by weighting the losses of the three proxy tasks. During training, the gradient of the loss function with respect to the model parameters is calculated through the back-propagation algorithm, and the model parameters are updated using the optimization algorithm.
10. The method for detecting leather anomalies with a small number of samples based on domain adversarial and multi-scale fusion according to claim 9, wherein: The model parameter optimization is divided into two stages: for the first agent task, the first stage comes from The images are used to optimize the model: ; in is the learning rate, is the parameter optimized by random alternation training of the first agent task, Represents the gradient of the loss function with respect to the parameters, which is used to indicate the optimization direction. The first agent task is to The loss function of training; in the second stage, similarly, using the target domain The image in the target domain is segmented by the loss Update model parameters ; are the updated model parameters; The same model parameter optimization method is used to implement the second and third proxy tasks.
Citation Information
Patent Citations
Cross-domain expression motion unit detection method based on text bridging
CN120145166A
Automatic unlabeled pancreas image segmentation system based on adversarial learning
WO2023098289A1