A fabric unsupervised anomaly detection method based on feature-pixel dual-domain collaborative reconstruction
Patent Information
- Application Number
- CN202610644845.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-12
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明要解决的技术问题是提供一种特征-像素双域协同重构的织物表面无监督异常检测方法,以解决现有无监督异常检测方法在重复纹理背景下对细微缺陷、方向扰动型缺陷和结构性异常感知不足的问题,从而提高织物表面缺陷的像素级定位精度与区域覆盖完整性
[0056](1)结构异常感知能力增强:通过在特征域中引入逐位置特征对齐与方向一致性约束,本发明不仅约束教师特征与学生特征在对应位置上的表征一致性,还进一步约束局部邻域内的方向结构关系,从而能够更有效地捕捉织物纹理中的断裂、扭曲、偏移和细长条纹类缺陷;
Smart Images

Figure CN122820536A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image detection and unsupervised anomaly detection technology, specifically relating to an unsupervised anomaly detection method for fabric surfaces based on feature-pixel dual-domain collaborative reconstruction. Background Technology
[0002] Fabric surface defect detection is a crucial step in quality control within the textile industry. In existing industrial settings, fabric images typically exhibit significant repetitive textures, and defect areas are often small, elongated, and highly similar in grayscale and texture to the normal background. This leads to inefficiencies, high subjectivity, and high false negative rates in traditional manual inspection methods. Existing deep learning-based fabric defect detection methods can be broadly categorized into supervised learning methods and unsupervised anomaly detection methods. Supervised learning methods rely on a large number of labeled defect samples, but in real-world industrial scenarios, defect samples are scarce and labeling is costly, hindering large-scale deployment. In contrast, unsupervised anomaly detection methods learn normal distributions using only normal samples, better meeting the needs of industrial applications.
[0003] However, existing unsupervised anomaly detection schemes still have the following shortcomings. First, methods based on feature distillation or feature distribution modeling typically focus on high-level semantic representation and lack sensitivity to fabric defects such as local texture direction changes and fine structural breaks. Second, while pixel-based reconstruction methods can generate pixel-level difference maps, their ability to recover local details lost during patch-level downsampling in the visual Transformer encoding process is limited, easily leading to blurred boundaries or discontinuous responses in anomalous regions. Third, existing schemes mostly model independently in the feature domain or pixel domain, lacking collaborative constraints on local structural relationships and texture details, resulting in frequent false positives and false negatives even against backgrounds with repetitive textures. Therefore, current technologies still lack an unsupervised anomaly detection method for fabric surfaces that can simultaneously achieve local structural consistency, texture detail recovery, and continuous coverage of anomalous regions. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an unsupervised anomaly detection method for fabric surface based on feature-pixel dual-domain collaborative reconstruction, so as to solve the problem that existing unsupervised anomaly detection methods are insufficient in perceiving subtle defects, directional perturbation defects and structural anomalies in the background of repetitive textures, thereby improving the pixel-level positioning accuracy and regional coverage integrity of fabric surface defects.
[0005] To address the aforementioned technical problems, this invention provides an unsupervised anomaly detection method for fabric surfaces based on feature-pixel dual-domain collaborative reconstruction, the specific process of which is as follows:
[0006] S1. Construct a feature-pixel dual-domain collaborative reconstruction network, which includes a student dual decoder composed of a teacher encoder, a feature domain reconstruction branch, and a pixel domain lightweight reconstruction branch. The original input image extracts multi-layer teacher features through the teacher encoder, and then outputs student reconstructed features through the feature domain reconstruction branch. The pixel domain lightweight reconstruction branch restores the resolution and texture details of the fabric image and outputs the reconstructed image.
[0007] S2. During the offline training phase, the positional feature alignment loss and orientation consistency loss are calculated based on the teacher features and the student reconstruction features; the pixel domain reconstruction loss is calculated based on the original input image and the reconstructed image; then, a joint loss function is constructed based on the positional feature alignment loss, orientation consistency loss and pixel domain reconstruction loss. During the training process, the teacher encoder parameters are kept frozen and the network parameters of the student dual decoder are iteratively optimized.
[0008] S3. In the online inference stage, images of the fabric to be detected are acquired in real time. After preprocessing, they are used as the original input images and input into the offline trained feature-pixel dual-domain collaborative reconstruction network. The feature domain reconstruction branch outputs student reconstructed features to generate position-wise feature alignment anomaly maps and orientation anomaly maps. The pixel domain lightweight reconstruction branch outputs reconstructed images to generate pixel domain anomaly maps. Then, the final anomaly map is obtained through a two-stage fusion strategy to achieve defect localization.
[0009] As an improvement to the unsupervised anomaly detection method for fabric surfaces based on feature-pixel dual-domain collaborative reconstruction of the present invention:
[0010] The teacher encoder uses a pre-trained visual Transformer model.
[0011] As a further improvement to the feature-pixel dual-domain collaborative reconstruction method for unsupervised anomaly detection on fabric surfaces of the present invention:
[0012] The feature domain reconstruction branch includes, in sequence, a bottleneck layer and a continuously stacked lightweight decoding block;
[0013] The pixel-domain lightweight reconstruction branch includes a total of four levels of pixel-domain decoder modules. The first three levels all use bilinear interpolation for upsampling operations, and then extract local texture information through two-layer depthwise separable convolution. The last level includes upsampling operations, convolutional layers, and a sigmoid activation function.
[0014] As a further improvement to the feature-pixel dual-domain collaborative reconstruction method for unsupervised anomaly detection on fabric surfaces of the present invention:
[0015] The position-wise feature alignment loss is calculated based on the responses of teacher features and student reconstructed features at spatial locations, and the formula is as follows:
[0016] (1)
[0017] in, This represents the position-wise feature alignment loss. The total number of feature layers. This represents the total number of spatial locations on a single-layer feature map. For the first Hierarchical normalization of teacher characteristics, For the first Hierarchical normalization of student reconstruction characteristics.
[0018] As a further improvement to the feature-pixel dual-domain collaborative reconstruction method for unsupervised anomaly detection on fabric surfaces of the present invention:
[0019] The directional consistency loss is calculated based on the difference between teacher features and student reconstructed features in spatial direction, and the formula is as follows:
[0020] (4)
[0021] in, This represents the cosine similarity between teacher features and student reconstructed features in a local direction.
[0022] (3)
[0023] in, To prevent extremely small constants with a denominator of zero, and These represent the joint gradient vector composed of the horizontal difference vector and the vertical difference vector, respectively.
[0024] (2)
[0025] in, and These represent the offsets between adjacent positions in the horizontal and vertical directions, respectively.
[0026] As a further improvement to the feature-pixel dual-domain collaborative reconstruction method for unsupervised anomaly detection on fabric surfaces of the present invention:
[0027] The pixel domain reconstruction loss is calculated based on the differences between the original input image and the reconstructed image in pixel values, structure, and image edge details, and the formula is as follows:
[0028] (5)
[0029] in, For mean square error loss, For structural similarity loss, For gradient consistency loss;
[0030] (6)
[0031] (7)
[0032] (8)
[0033] in, and These represent the gradient operators in the horizontal and vertical directions, respectively. and These represent the height and width of the image, respectively. Represents pixel coordinates, Represents the original input image. Indicates the reconstructed image. This represents the structural similarity index.
[0034] As a further improvement to the feature-pixel dual-domain collaborative reconstruction method for unsupervised anomaly detection on fabric surfaces of the present invention:
[0035] The joint loss function is:
[0036] (9)
[0037] in, , and These are the weight coefficients for position-wise feature alignment loss, orientation consistency loss, and pixel domain reconstruction loss, respectively.
[0038] As a further improvement to the feature-pixel dual-domain collaborative reconstruction method for unsupervised anomaly detection on fabric surfaces of the present invention:
[0039] The position-by-position feature alignment anomaly map is defined as follows:
[0040] (10)
[0041] The directional anomaly map is defined as follows:
[0042] (11)
[0043] The pixel domain anomaly map is defined as follows:
[0044] (12).
[0045] As a further improvement to the feature-pixel dual-domain collaborative reconstruction method for unsupervised anomaly detection on fabric surfaces of the present invention:
[0046] The two-stage fusion strategy includes:
[0047] The anomaly map obtained after the first-stage fusion is calculated using the following formula:
[0048] (13)
[0049] in, This indicates min-max normalization. These are the weight parameters for the directional anomaly map;
[0050] The final anomaly map is obtained after the second stage of fusion:
[0051] (14)
[0052] in, and These are the resampled and normalized comprehensive feature domain anomaly map and pixel domain anomaly map, respectively. The scaling parameter for the pixel domain anomaly map is 'max', where 'max' indicates the scaling parameter calculated by the maximum value of each element.
[0053] As a further improvement to the feature-pixel dual-domain collaborative reconstruction method for unsupervised anomaly detection on fabric surfaces of the present invention:
[0054] During the offline training phase, the feature-pixel dual-domain collaborative reconstruction network uses a training set consisting only of normal samples and a test set consisting of both normal and abnormal samples.
[0055] The beneficial effects of this invention are mainly reflected in:
[0056] (1) Enhanced ability to perceive structural anomalies: By introducing positional feature alignment and orientation consistency constraints in the feature domain, this invention not only constrains the representational consistency of teacher features and student features at corresponding positions, but also further constrains the orientation structure relationship in the local neighborhood, thereby enabling more effective capture of defects such as breaks, twists, offsets and thin stripes in fabric textures.
[0057] (2) Enhanced texture detail recovery capability: By introducing a pixel domain lightweight reconstruction branch with a low parameter amount, this invention can compensate for the loss of low-level texture information caused by patch-level representation compression during the visual Transformer encoding process, thereby improving the problems of blurred boundaries, discrete response and incomplete coverage in abnormal regions;
[0058] (3) Dual-domain collaborative localization is more stable: This invention adopts dual-domain collaborative modeling of feature domain and pixel domain, and complements and enhances structural anomaly response and texture difference response through a two-stage fusion strategy, which effectively suppresses background false alarms while improving the coverage integrity of anomaly areas.
[0059] (4) Low parameter overhead and real-time deployment capability: The pixel domain reconstruction branch adopts a lightweight convolutional structure, which improves the pixel-level defect localization performance while introducing only a small amount of additional parameters, taking into account both detection accuracy and inference efficiency, and is suitable for industrial online detection scenarios.
[0060] (5) Applicable to unsupervised industrial inspection scenarios: This invention relies only on normal samples for training and does not require a large amount of defect labeling data, which can better adapt to the application conditions of rare abnormal samples and uneven category distribution in real industrial production. Attached Figure Description
[0061] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0062] Figure 1 This is a schematic diagram of the overall structure of the feature-pixel dual-domain collaborative reconstruction unsupervised anomaly detection framework of the present invention;
[0063] Figure 2 This is a schematic diagram of the feature domain lightweight decoding structure of the present invention;
[0064] Figure 3 This is a schematic diagram of the pixel domain lightweight reconstruction branch structure of the present invention;
[0065] Figure 4 This is a schematic diagram of a typical sample from the Fabric dataset of real industrial fabrics in this invention;
[0066] Figure 5 This is a visual comparison of the defect detection results of the present invention and the comparison method on the Fabric dataset. Detailed Implementation
[0067] The present invention will be further described below with reference to specific embodiments, but the scope of protection of the present invention is not limited thereto:
[0068] Example 1: An unsupervised method for fabric defect detection based on feature-pixel dual-domain collaborative reconstruction, such as... Figure 1 As shown, the process includes teacher feature extraction, dual-domain collaborative reconstruction, joint loss optimization, and anomaly graph fusion inference stages. The specific process is as follows:
[0069] 1. Construct a feature-pixel dual-domain collaborative reconstruction network
[0070] like Figure 1As shown, this invention is based on a multi-class unsupervised anomaly detection model (Dinomaly model) and constructs a feature-pixel dual-domain collaborative reconstruction network based on a teacher-student distillation paradigm. Addressing the limitation of the basic model being limited to feature domain reconstruction, resulting in limited recovery of local fabric details and insufficient perception of subtle structural orientations, at the network topology level, this invention breaks through the limitations of a single decoder and constructs a student dual-decoder structure composed of a feature domain reconstruction branch and a newly added pixel domain lightweight reconstruction branch. This compensates for the loss of fine-grained texture information caused by deep feature compression. At the objective constraint level, this invention introduces position-wise feature alignment constraints and orientation consistency constraints in the feature domain and performs collaborative optimization in conjunction with the multi-dimensional reconstruction loss of the pixel domain. The feature-pixel dual-domain collaborative reconstruction network includes a frozen teacher encoder and a student dual decoder. The teacher encoder is used to extract multi-layer semantic features from the original input image and provide a stable normal distribution representation. The feature domain reconstruction branch of the student dual decoder is used to reconstruct teacher-side features to learn the distribution pattern of normal texture in the deep representation space; the pixel domain lightweight reconstruction branch is used to recover local texture details at the original image resolution to compensate for the insufficient recovery of low-level details by deep feature modeling. The two branches of the student dual decoder share the representation of the teacher encoder output and are jointly optimized through joint loss.
[0071] 1.1 Constructing a teacher encoder
[0072] The input fabric image is denoted as: the original input image. In this embodiment, the model uses a pre-trained visual Transformer as the teacher encoder, that is, a pre-trained DINOv2 model, such as the ViT-B / 14-Reg model. In this embodiment, the original input image... First, the tokens are divided into fixed-size Patch Tokens and input into a pre-trained visual Transformer teacher encoder to extract token features from multiple specified layers. To facilitate subsequent feature reconstruction and spatial alignment, the Patch Tokens from each selected layer, after removing the aggregated tokens, are rearranged into a two-dimensional feature map, resulting in the... Characteristics of teachers at different levels: Among them, T, B, and C l H l W l These represent the teacher, batch size, number of channels, and image height and width, respectively. The multi-layered teacher features output by the teacher encoder are used as shared features input to the student dual decoder.
[0073] 1.2 Constructing the Feature Domain Reconstruction Branch and its Loss Function
[0074] 1.2.1 Constructing the Feature Domain Reconstruction Branch
[0075] The feature domain reconstruction branch aims to recover deep feature representations from the semantic information extracted by the teacher encoder. In this embodiment, the feature domain reconstruction branch includes a bottleneck layer and eight consecutively stacked lightweight decoding blocks based on Transformer. The bottleneck layer injects noise into the input teacher features and compresses them; the output bottleneck features serve as the input to the lightweight decoding blocks. The eight consecutively stacked lightweight decoding blocks, through layer-by-layer feature transformation and relation modeling, ultimately output student reconstructed features corresponding to the teacher's side. This design enables the model to effectively capture and recover deep semantic information with low parameter overhead. Each lightweight decoding block employs a rigorous residual connection structure, such as... Figure 2 As shown, the feature flow sequentially passes through the first layer normalization, a linear attention mechanism, the first residual addition, the second layer normalization, a feedforward network (MLP) with a Gaussian Error Linear Unit (GELU) activation function, and the second residual addition. Through this structure, the lightweight decoding block can achieve feature relation modeling and representation reconstruction while maintaining low computational complexity. The output of the feature domain reconstruction branch is the first... The reconstruction characteristics of the students in the layer are: , where S represents student.
[0076] 1.2.2 Position-by-position feature alignment loss
[0077] To improve the comparability between features at different levels, the teacher features and the reconstructed student features are normalized in the channel dimension before calculating the loss, denoted as the _i_th. Hierarchical Normalization of Teacher Characteristics and the Hierarchical Normalization Student Reconstruction Features .
[0078] This invention first employs a position-by-position feature alignment method to constrain the consistency of responses between teacher features and student reconstructed features at the same spatial location, with the loss function being:
[0079] (1)
[0080] in, This represents the position-wise feature alignment loss. The total number of feature layers participating in the reconstruction. This represents the total number of spatial locations on a single-layer feature map. This serves as a spatial location index. The aforementioned loss can directly measure the consistency of the location of normal samples in the teacher's feature space and the student's reconstructed feature space.
[0081] 1.2.3 Loss of Directional Consistency
[0082] When using only position-wise feature alignment loss, the model mainly focuses on the representational differences at the same spatial location, but it is difficult to explicitly model the directional continuity and structural relationship of fabric texture in the local neighborhood. Considering that fabric texture has significant periodicity and directional features, this invention further introduces directional consistency constraints, and calculates the directional consistency loss based on the difference between teacher features and student reconstructed features in spatial direction.
[0083] For the Location in layer feature map Calculate the forward difference of teacher characteristics in the horizontal direction respectively. Forward difference of teacher characteristics in the vertical direction Forward difference of student reconstructed features in the horizontal direction Student-reconstructed features in the vertical direction forward difference :
[0084] (2)
[0085] in, and These represent the offsets of adjacent positions in the horizontal and vertical directions, respectively. When calculating the forward difference, to avoid edge positions going out of bounds, the horizontal gradient is calculated only in positions other than the rightmost column, and the vertical gradient is calculated only in positions other than the bottommost row. Then, zeros are padded at the rightmost side of the horizontal gradient map and at the bottommost side of the vertical gradient map, respectively, to keep the gradient map size consistent with the original feature map.
[0086] Furthermore, cosine similarity is used to measure the consistency of teacher features and student reconstructed features in local directional structure:
[0087] (3)
[0088] in, To prevent extremely small constants with a denominator of zero, and These represent the joint gradient vector composed of the horizontal difference vector and the vertical difference vector, respectively.
[0089] The directional consistency loss can be obtained by averaging across all locations and layers. :
[0090] (4)
[0091] This constraint requires the model to reconstruct not only single-point feature values but also local orientation structures when reconstructing normal textures, thereby improving the model's ability to respond to texture orientation mutations.
[0092] 1.3 Constructing the pixel-domain lightweight reconstruction branch and its loss function
[0093] The pixel-domain lightweight reconstruction branch is used to recover the original input image. The fine texture information at the original resolution is represented by the output reconstructed image, denoted as... The pixel-domain lightweight reconstruction branch adopts a step-by-step upsampling decoding structure, which uses a total of four levels of pixel-domain decoder modules, such as... Figure 3 As shown, the first three pixel-domain decoder modules each consist of upsampling operations and two-layer depthwise separable convolutions. Each pixel-domain decoder module first uses bilinear interpolation for upsampling to restore scale, and then uses two-layer depthwise separable convolutions to extract local texture information and reduce the number of parameters. The depthwise separable convolutions are composed of… Depth convolution and It is constructed using pointwise convolution. This design avoids the redundant computation caused by using large-scale ordinary convolution, restoring the original resolution texture information while maintaining a low number of parameters. This is an important component of this invention for improving pixel-level positioning accuracy. The final stage of the pixel-domain lightweight reconstruction branch consists of upsampling operations, It consists of convolution and the Sigmoid activation function.
[0094] Based on the original input image With reconstructed images To account for differences in pixel values, structure, and image edge details, this invention employs the following pixel-domain reconstruction loss function:
[0095] (5)
[0096] in, For mean square error loss, For structural similarity loss, This is the gradient consistency loss.
[0097] The three losses are specifically defined as follows:
[0098] (6)
[0099] (7)
[0100] (8)
[0101] in, and These represent the gradient operators in the horizontal and vertical directions, respectively. and These represent the height and width of the image, respectively. Indicates the first in the image line, number Pixel coordinates at column position Represents the original input image With reconstructed images The structural similarity index between them.
[0102] 2. Offline model training
[0103] 2.1 Dataset Creation
[0104] The dataset used in this invention is the Fabric dataset, a real-world industrial fabric dataset. This dataset consists of fabric surface images collected from actual industrial production scenarios, including both normal and abnormal fabric samples. Abnormality types include weft insertion, warp breakage, weft breakage, frilly weft, and loose weft. The collected fabric images are preprocessed (including cropping, size standardization, and normalization) before being used as the original input images. Input the model for training and testing. Some normal and abnormal samples from the Fabric dataset are shown below. Figure 4 As shown.
[0105] To meet the training requirements of unsupervised anomaly detection tasks, the dataset is divided as follows: the training set contains only images of normal fabrics to learn the normal texture distribution; the test set includes both normal and anomaly samples to evaluate the model's image-level anomaly detection capability and pixel-level defect localization capability. In this embodiment, the Fabric dataset includes 252 normal samples in the training set, 100 normal samples in the test set, and 1107 anomaly samples in the test set.
[0106] 2.2 Design of Joint Loss Function
[0107] To simultaneously learn the distribution patterns of normal fabric in both the feature space and pixel space, this invention employs the following joint loss function to jointly optimize the dual-domain collaborative reconstruction network:
[0108] (9)
[0109] in, , and These are the weight coefficients for position-wise feature alignment loss, orientation consistency loss, and pixel domain reconstruction loss, respectively.
[0110] In this embodiment, , and It can be set to 1.0, 1.0, and 1.0; it can also be adjusted based on the validation set performance under different fabric texture types or image resolutions. The value can be appropriately increased to enhance the model's sensitivity to defects caused by texture direction perturbations. It can be appropriately increased to enhance the ability to reconstruct pixel-domain details.
[0111] 2.3 Offline Training and Testing Process
[0112] The dataset established in step 2.1 is input into the constructed dual-domain collaborative reconstruction network for training. During the training phase, only normal fabric samples are used. Multi-layer features are extracted through the teacher encoder, and feature reconstruction and image reconstruction are completed by the feature domain reconstruction branch and pixel domain lightweight reconstruction branch in the student network, respectively.
[0113] In this embodiment, the teacher encoder parameters are kept frozen, and only the student network parameters are updated. The AdamW optimizer can be used during training, and the learning rate can be set to... The optimal batch size is 16, and the optimal number of training epochs is 1000. During training, optimization is performed based on the joint loss function shown in step 2.2. The student network parameters are iteratively optimized using the backpropagation algorithm, and the model weights with the best training effect are saved.
[0114] During the testing phase, test set images were input into the trained network to generate image-level anomaly scores and pixel-level anomaly maps. The anomaly detection performance of the model was then evaluated based on evaluation metrics. In the image-level anomaly detection task, I-AUROC and I-AP reached 99.76% and 99.95%, respectively. In the pixel-level localization task, P-AUROC achieved 94.13% and P-PRO achieved 84.53%, thus obtaining a dual-domain collaborative reconstruction network that can be used online to achieve defect localization.
[0115] 3. Online use
[0116] like Figure 1 As shown, during the online inspection phase, images of the completed fabric, acquired in real time, are preprocessed (including cutting, size standardization, and normalization) and used as the original input images. The input is fed into an offline-trained feature-pixel dual-domain collaborative reconstruction network, and the teacher encoder extracts the original input image. Multi-layered teacher features are used as shared feature inputs to generate dual-domain anomaly maps using a student dual decoder: the feature domain reconstruction branch outputs student reconstructed features to generate position-wise feature-aligned anomaly maps and orientation anomaly maps, and the pixel domain lightweight reconstruction branch outputs the reconstructed image. This is used to generate pixel-domain anomaly maps, and then a two-stage fusion strategy is used to obtain the final anomaly map.
[0117] The per-position feature alignment anomaly map is defined as:
[0118] (10)
[0119] The directional anomaly map is defined as:
[0120] (11)
[0121] Pixel domain anomaly map is defined as:
[0122] (12)
[0123] Considering that the three types of anomaly graphs reflect anomaly information at different levels: It mainly reflects differences in characteristic representation. It mainly reflects changes in local structural direction. Since it mainly reflects differences in texture details, this invention adopts a two-stage fusion strategy to generate the final anomaly map.
[0124] The first stage involves normalizing the position-wise feature alignment anomaly map and the orientation anomaly map, and then performing linear fusion according to weights to obtain the comprehensive feature domain anomaly map. :
[0125] (13)
[0126] in, This indicates min-max normalization. These are the weight parameters for the directional anomaly map. In this embodiment, The preferred value range is 0.3 to 0.7. and All are calculated from the positional correspondence in the same visual Transformer feature space, and their spatial resolution is consistent, so they can be directly fused element by element.
[0127] The second stage involves further fusing the comprehensive feature domain anomaly map with the pixel domain anomaly map to obtain the final anomaly map. This is because the comprehensive feature domain anomaly map... Based on feature space location calculation, its spatial resolution is affected by the feature map downsampling rate; pixel domain anomaly map The spatial resolution is calculated from the difference between the reconstructed image and the original input image, and it depends on the output size of the reconstruction branch. Therefore, before performing cross-domain fusion, the spatial resolution is first calculated from the difference between the reconstructed image and the original input image. and The image is resampled to the original input image size using bilinear interpolation. For simplicity, the resampled composite feature domain anomaly map and pixel domain anomaly map are then normalized and denoted as follows: and After scaling proportionally, the final anomaly graph is obtained by calculating the maximum value of each element.
[0128] (14)
[0129] Here, max represents the operation by maximizing the value of each element. This is a scaling parameter for the pixel-domain anomaly map, used to adjust the contribution ratio of pixel-domain responses in the final anomaly map. In this embodiment, The preferred value range is 0.5 to 0.8. The two-stage fusion method first integrates structural anomaly information within the feature domain, and then combines it with pixel domain texture difference information, which can effectively improve the coverage integrity of the anomaly area and suppress false background alarms.
[0130] 4. Experiment
[0131] 4.1 Evaluation Indicators
[0132] To verify the detection performance of the method of this invention on the real industrial fabric dataset Fabric, this embodiment uses image-level evaluation metrics and pixel-level evaluation metrics to comprehensively evaluate the model. Image-level evaluation metrics include I-AUROC and I-AP, where I-AUROC measures the model's overall ability to distinguish between normal and abnormal samples, and I-AP measures the model's anomaly detection performance under imbalanced sample conditions. Pixel-level evaluation metrics include P-AUROC and P-PRO, where P-AUROC measures the overall separability between abnormal and normal pixels, and P-PRO measures the overall coverage of abnormal regions, which is more suitable for evaluating the performance of fabric surface defect localization.
[0133] 4.2 Comparative Experiment
[0134] To verify the effectiveness of the method of this invention, six representative unsupervised anomaly detection methods were selected as comparison methods, including PatchCore, RD, FastFlow, SuperSimpleNet, EfficientAD, and Dinomaly. The performance metrics comparison results are shown in Table 1, and some visualization comparisons of the detection results are shown below. Figure 5 As shown.
[0135] As shown in Table 1, the method of the present invention maintains a high level of anomaly discrimination capability in image-level anomaly detection tasks, with I-AUROC and I-AP reaching 99.76% and 99.95%, respectively. Although slightly lower than Dinomaly and EfficientAD in image-level metrics, the difference is small, indicating that the method of the present invention does not weaken the overall anomaly recognition capability due to the introduction of the dual-domain collaborative reconstruction mechanism.
[0136] In pixel-level localization tasks, the method of this invention achieved a P-AUROC of 94.13% and a P-PRO of 84.53%, both outperforming all comparative methods. Specifically, compared to Dinomaly, P-AUROC improved by 2.16 percentage points, and P-PRO by 6.95 percentage points; compared to PatchCore, P-PRO improved by 2.91 percentage points. These results demonstrate that the method of this invention can more completely cover abnormal areas in real industrial fabric scenarios and effectively improve the localization capability of small defects and structural disturbance defects.
[0137] Table 1. Detection results of each method on the Fabric dataset.
[0138]
[0139] Depend on Figure 5 As can be seen, compared with the comparative method, the method of the present invention can obtain a more continuous and complete abnormal response area for small defects, elongated defects and texture direction disturbance defects, indicating that the method of the present invention has advantages in terms of the completeness of abnormal area coverage.
[0140] 4.3 Ablation Experiment
[0141] Table 2 presents the results of ablation experiments on the Fabric dataset, validating the contribution of each module. Using the Dinomaly model as the base model (first row of Table 2), the effects of modules A (pixel-domain lightweight reconstruction branch), B (position-wise feature alignment constraint), and C (directional consistency constraint) are examined. Module A provides the most direct improvement to pixel-level localization. After introducing module A, the feature-pixel dual-domain collaborative reconstruction network of this invention increases from 77.58% to 82.24% of the baseline, indicating that the pixel-domain branch effectively supplements local texture and improves the ability to locate subtle defects. Module B strengthens the local correspondence between teacher features and student reconstructed features through stricter point-to-point feature constraints; Module C, starting from local directional relationships, enhances the model's sensitivity to texture direction perturbations and structural anomalies. As can be seen from the table, the two types of constraints significantly improve the feature-pixel dual-domain collaborative reconstruction network of this invention. Meanwhile, when module A is combined with only a single feature constraint (either B or C alone), the P-PRO metric shows a slight decrease compared to the single feature constraint. This indicates that introducing pixel reconstruction can lead to optimization conflicts when the feature domain structure constraints are incomplete. However, when all three modules are used in conjunction, the model achieves optimal performance on the dataset. This demonstrates a clear synergy between the lightweight pixel domain reconstruction branch and the complete feature domain constraints: the former focuses on restoring local texture, while the latter focuses on enhancing the response to anomalous structural relationships and directional changes. The combination of the two can more stably improve the coverage integrity and localization accuracy of defective regions.
[0142] Table 2 Ablation experiment results of this method on the Fabric dataset (%)
[0143]
[0144] Regarding model complexity, because the lightweight pixel-domain reconstruction branch in this invention employs a lightweight depthwise separable convolution design, the number of parameters in this invention only increases from approximately 61.4M to approximately 61.7M compared to the baseline Dinomaly, with an additional parameter overhead of approximately 0.3M. In terms of inference efficiency, under the standard test environment of an RTX 4090 (24GB) graphics processing unit (GPU), this invention achieves 79.52 FPS (Frames Per Second), demonstrating that while significantly improving pixel-level localization performance, this invention still possesses good feasibility for industrial online detection.
[0145] In summary, the method of this invention demonstrates excellent unsupervised anomaly detection capabilities on the real industrial fabric dataset Fabric. Compared with existing methods, this invention achieves a significant improvement in pixel-level localization metrics (P-PRO increases by 6.95%) by slightly sacrificing image-level metrics (I-AUROC decreases slightly from 100% to 99.76%). While maintaining a high level of image-level detection performance, it further significantly improves pixel-level localization accuracy and defect region coverage integrity, verifying the effectiveness of the feature domain and pixel domain collaborative reconstruction mechanism and its application value in industrial fabric defect detection scenarios.
[0146] Finally, it should be noted that the above examples are merely some specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments and many variations are possible. All variations that can be directly derived or conceived by those skilled in the art from the disclosure of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A method for unsupervised anomaly detection on fabric surfaces using feature-pixel dual-domain collaborative reconstruction, characterized in that... The specific process includes the following: S1. Construct a feature-pixel dual-domain collaborative reconstruction network, which includes a student dual decoder composed of a teacher encoder, a feature domain reconstruction branch, and a pixel domain lightweight reconstruction branch. The original input image extracts multi-layer teacher features through the teacher encoder, and then outputs student reconstructed features through the feature domain reconstruction branch. The pixel domain lightweight reconstruction branch restores the resolution and texture details of the fabric image and outputs the reconstructed image. S2. During the offline training phase, the positional feature alignment loss and orientation consistency loss are calculated based on the teacher features and the student reconstruction features; the pixel domain reconstruction loss is calculated based on the original input image and the reconstructed image; then, a joint loss function is constructed based on the positional feature alignment loss, orientation consistency loss and pixel domain reconstruction loss. During the training process, the teacher encoder parameters are kept frozen and the network parameters of the student dual decoder are iteratively optimized. S3. In the online inference stage, images of the fabric to be detected are acquired in real time. After preprocessing, they are used as the original input images and input into the offline trained feature-pixel dual-domain collaborative reconstruction network. The feature domain reconstruction branch outputs student reconstructed features to generate position-wise feature alignment anomaly maps and orientation anomaly maps. The pixel domain lightweight reconstruction branch outputs reconstructed images to generate pixel domain anomaly maps. Then, the final anomaly map is obtained through a two-stage fusion strategy to achieve defect localization.
2. The unsupervised anomaly detection method for fabric surface based on feature-pixel dual-domain collaborative reconstruction according to claim 1, characterized in that: The teacher encoder uses a pre-trained visual Transformer model.
3. The unsupervised anomaly detection method for fabric surface based on feature-pixel dual-domain collaborative reconstruction according to claim 2, characterized in that: The feature domain reconstruction branch includes, in sequence, a bottleneck layer and a continuously stacked lightweight decoding block; The pixel-domain lightweight reconstruction branch includes a total of four levels of pixel-domain decoder modules. The first three levels all use bilinear interpolation for upsampling operations, and then extract local texture information through two-layer depthwise separable convolution. The last level includes upsampling operations, convolutional layers, and a sigmoid activation function.
4. The unsupervised anomaly detection method for fabric surface based on feature-pixel dual-domain collaborative reconstruction according to claim 3, characterized in that: The position-wise feature alignment loss is calculated based on the responses of teacher features and student reconstructed features at spatial locations, and the formula is as follows: (1) in, This represents the position-wise feature alignment loss. The total number of feature layers. This represents the total number of spatial locations on a single-layer feature map. For the first Hierarchical normalization of teacher characteristics, For the first Hierarchical normalization of student reconstruction characteristics.
5. The unsupervised anomaly detection method for fabric surface based on feature-pixel dual-domain collaborative reconstruction according to claim 4, characterized in that: The directional consistency loss is calculated based on the difference between teacher features and student reconstructed features in spatial direction, and the formula is as follows: (4) in, This represents the cosine similarity between teacher features and student reconstructed features in a local direction. (3) in, To prevent extremely small constants with a denominator of zero, and These represent the joint gradient vector composed of the horizontal difference vector and the vertical difference vector, respectively. (2) in, and These represent the offsets between adjacent positions in the horizontal and vertical directions, respectively.
6. The unsupervised anomaly detection method for fabric surface based on feature-pixel dual-domain collaborative reconstruction according to claim 5, characterized in that: The pixel domain reconstruction loss is calculated based on the differences between the original input image and the reconstructed image in pixel values, structure, and image edge details, and the formula is as follows: (5) in, For mean square error loss, For structural similarity loss, For gradient consistency loss; (6) (7) (8) in, and These represent the gradient operators in the horizontal and vertical directions, respectively. and These represent the height and width of the image, respectively. Represents pixel coordinates, Represents the original input image. Indicates the reconstructed image. This represents the structural similarity index.
7. The unsupervised anomaly detection method for fabric surface based on feature-pixel dual-domain collaborative reconstruction according to claim 6, characterized in that: The joint loss function is: (9) in, , and These are the weight coefficients for position-wise feature alignment loss, orientation consistency loss, and pixel domain reconstruction loss, respectively.
8. The unsupervised anomaly detection method for fabric surface based on feature-pixel dual-domain collaborative reconstruction according to claim 7, characterized in that: The position-by-position feature alignment anomaly map is defined as follows: (10) The directional anomaly map is defined as follows: (11) The pixel domain anomaly map is defined as follows: (12)。 9. The unsupervised anomaly detection method for fabric surface based on feature-pixel dual-domain collaborative reconstruction according to claim 8, characterized in that: The two-stage fusion strategy includes: The anomaly map obtained after the first-stage fusion is calculated using the following formula: (13) in, This indicates min-max normalization. These are the weight parameters for the directional anomaly map; The final anomaly map is obtained after the second stage of fusion: (14) in, and These are the resampled and normalized comprehensive feature domain anomaly map and pixel domain anomaly map, respectively. The scaling parameter for the pixel domain anomaly map is 'max', where 'max' indicates the scaling parameter calculated by the maximum value of each element.
10. The unsupervised anomaly detection method for fabric surface based on feature-pixel dual-domain collaborative reconstruction according to claim 9, characterized in that: During the offline training phase, the feature-pixel dual-domain collaborative reconstruction network uses a training set consisting only of normal samples and a test set consisting of both normal and abnormal samples.