Industrial image anomaly detection method based on self-supervised flow model
The D3Sflow model is constructed by combining the self-supervised paradigm and multi-scale dynamic feature fusion technology, which solves the problems of low discrimination of abnormal feature distribution mapping results and insufficient modeling capabilities of complex multi-scale features in existing technologies, and achieves high precision and robustness in industrial image anomaly detection.
Patent Information
- Application Number
- CN202510758504.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-10-10
AI Technical Summary
Existing industrial image anomaly detection methods based on normalized flow have problems such as low discrimination of anomaly feature distribution mapping results and insufficient ability to model complex multi-scale feature distributions, resulting in insufficient detection accuracy.
The self-supervised paradigm and multi-scale dynamic feature fusion technology are adopted to construct the D3Sflow model. Through feature extraction, preprocessing, and normalized flow distribution mapping modules, the pseudo-anomaly generation strategy and collaborative likelihood optimization loss function are used to optimize the model parameters and improve the discriminability of feature mapping and multi-scale modeling capabilities.
It significantly improves the accuracy and robustness of anomaly detection in industrial scenarios, can more accurately fit the normal feature distribution, improves the ability to distinguish positive and abnormal features, and reduces the model calculation amount.
Smart Images

Figure CN120765987A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial image processing, and in particular relates to an industrial image anomaly detection method based on a self-supervised flow model. Background Art
[0002] Industrial image anomaly detection is a key technology in intelligent manufacturing. Its goal is to automatically identify cosmetic defects or structural anomalies in industrial products using computer vision techniques. Depending on whether the training samples contain abnormal samples, anomaly detection models can be divided into supervised and unsupervised settings. Supervised models require a large number of accurately labeled abnormal images for training. However, in real industrial scenarios, abnormal samples are scarce and labeling is expensive, so unsupervised anomaly detection methods are gradually becoming mainstream. Among them, models based on Normalizing Flow (NF) demonstrate high detection efficiency and accuracy by mapping normal features to a multivariate standard Gaussian distribution and detecting anomalies based on the likelihood of the mapping result.
[0003] However, existing industrial anomaly detection methods based on normalized flows have the following drawbacks:
[0004] (1) The mapping results of abnormal feature distributions are not distinguishable: In unsupervised training, the model only learns the distribution mapping of normal features and lacks explicit modeling of abnormal features. This causes the mapping results of abnormal features to overlap with normal features, reducing the model's ability to distinguish unknown anomalies.
[0005] (2) Insufficient modeling capabilities for complex multi-scale feature distributions: Industrial products have complex structures and significant differences in defect scales (e.g., tiny scratches versus large-scale dislocations). Existing normalized flow models are not structurally optimized for modeling complex multi-scale feature distributions, resulting in limited comprehensive detection capabilities for complex textures and large-scale anomalies. Summary of the Invention
[0006] The purpose of this invention is to provide an industrial image anomaly detection method based on a self-supervised flow model. This method aims to address the existing problems of insufficient detection accuracy due to the low discrimination of anomaly feature distribution mapping results and the lack of complex multi-scale feature modeling capabilities. By integrating a self-supervised paradigm with multi-scale dynamic feature fusion technology, this method significantly improves the accuracy and robustness of anomaly detection in industrial scenarios.
[0007] The present invention provides an industrial image anomaly detection method based on a self-supervised flow model. The method includes constructing a dynamic selective self-supervised flow (D3Sflow) model with dynamic feature selection. The model includes a feature extraction module, a feature preprocessing module, and a normalized flow (NF) distribution mapping module. The model is then trained and tested. The specific steps are as follows:
[0008] S1. Input the training image into the feature extraction module and the feature preprocessing module in sequence to obtain feature maps at three different scales containing normal features and pseudo-abnormal features (i.e., multi-scale positive and abnormal feature maps);
[0009] S2. Input the multi-scale positive and abnormal feature map obtained in S1 into the NF distribution mapping module, mapping the input features in the original feature space (distribution unknown) to the multivariate standard Gaussian distribution in the latent space (distribution known), so that the likelihood value of the input features after mapping by this module can distinguish positive and abnormal features;
[0010] S3. Update model parameters by minimizing the self-supervised loss function based on collaborative likelihood optimization, optimize the distribution of normal features and pseudo-anomaly features, and gradually approach the overall optimization goal of the D3SFlow model;
[0011] S4. Input the image to be detected into the trained D3SFlow model, calculate the anomaly score of the model output, and generate image-level and pixel-level anomaly detection results.
[0012] Among them, S1-S3 are the training stage, and S4 is the detection stage.
[0013] Furthermore, in step S1, the feature extraction module can select two different structures according to the different dimensions of the pseudo-anomaly generation strategy used: one is the module structure under the image dimension, first obtaining normal and pseudo-anomaly images through the image dimension pseudo-anomaly generation step, and then performing pre-training feature extraction to obtain multi-scale positive anomaly features; the other is the module structure under the feature dimension, first obtaining multi-scale normal features through pre-training feature extraction, and then obtaining pseudo-anomaly features through the feature dimension pseudo-anomaly generation step.
[0014] The image dimension pseudo anomaly generation process can be divided into the following steps:
[0015] (1) Generate pseudo-abnormal region masks through foreground-enhanced Perlin noise mask generator;
[0016] (2) Generate images as anomaly sources through structural and texture anomaly generators, generate pseudo structural anomaly images from the same source dataset with probability p, and generate pseudo texture anomaly images from the heterogeneous dataset with probability 1-p;
[0017] (3) First, the abnormal area mask generated by the foreground enhanced Perlin noise mask generator is converted into a normal area mask by inverting the pixel values, and then the final simulated abnormal image is generated using the following formula:
[0018]
[0019] Among them, M is the pseudo abnormal area mask generated by the foreground enhanced Perlin noise mask generator, M is the normal area mask generated by pixel value inversion, N is the normal image to be generated as a pseudo-anomaly, A is the abnormal source image generated by the structure and texture anomaly generator, and β is the abnormality degree factor with a value interval of (0.2, 0.8), which means that the abnormal source is added to the abnormal area at a ratio of β and the original image is retained at a ratio of 1-β.
[0020] The feature dimension pseudo-anomaly generation process can be divided into the following steps:
[0021] (1) Generate abnormal area masks (M1, M2, M3) corresponding to feature maps of different sizes through the pyramid mask generator;
[0022] (2) Through the feature perturbation of foreground enhancement, the corresponding masks (M1, M2, M3) are used to add noise perturbation to the feature maps (h1, h2, h3) of different scales according to the following formula to obtain the corresponding pseudo-abnormal feature h A , where ∈ is independently sampled from a Gaussian distribution
[0023] h A =h N +M i ⊙∈,
[0024] Among them, h N It is one of the normal feature maps (h1, h2, h3) processed by the feature perturbation of foreground enhancement, M i is the abnormal area mask (M1, M2, M3) generated by the pyramid mask generator and h N Corresponding masks of the same size;
[0025] (3) Concatenate normal features and pseudo-abnormal features in the batch dimension and output a multi-scale feature map containing both normal features and pseudo-abnormal features (where the batch size is twice the size of the input feature map).
[0026] In the pre-training feature extraction process, the dataset is divided into a training set and a test set according to the requirements of the single-category industrial anomaly detection (OC-IAD) task. The images in the union of the two sets are input into a CNN pre-trained network Φ with fixed parameters, and the last feature map of multiple stages of the network is sampled. where L=3 means that the feature maps of the first three stages are extracted, so the pre-training feature extraction process can be represented as the following formula:
[0027] (h1,h2,h3)=Φ(x),
[0028] where h1, h2, h3 represent the feature maps of three different scales respectively.
[0029] Further, the feature preprocessing module in step S1 has the following workflow:
[0030] (1) The input feature map is down-sampled by a learnable down-sampling module. Specifically, the input feature is first decomposed into a diagonal direction high-frequency component y D , a vertical direction high-frequency component y V , a horizontal direction high-frequency component y H , and a low-frequency component y A by a two-dimensional Haar wavelet transform module, and then the channel dimension is spliced, and the channel number is aligned by a learnable feature alignment module.
[0031] (2) The pre-trained feature after down-sampling is mapped to the target domain using a feature adapter module.
[0032] Further, in step S2, the NF distribution mapping module has the following workflow:
[0033] (1) Single-scale feature modeling, which is mainly responsible for learning the internal attributes of a single scale feature map, i.e., learning the mapping method that makes the likelihood of each scale feature map maximum by three parallel NFs;
[0034] (2) Multi-scale feature modeling, which is mainly responsible for sharing the global semantic information and local detail information contained in different scale feature maps, i.e., learning the mapping method that makes the comprehensive likelihood of the input feature map under multiple scales maximum by a single NF, so that the model can simultaneously utilize different effective information in multi-scale feature maps for anomaly detection.
[0035] The single-scale feature modeling stage uses a deformable convolution attention parallel flow (DAPF) module to independently model the feature maps of three scales. In order to improve the distribution mapping efficiency and reduce the overall computational amount of the model, a pyramid-like structure is used, and different numbers (proportional to the number of channels of the feature) of DAPF modules (containing an affine coupling layer of continuous two opposite transformation processes) are used to map the distributions of different scale features.
[0036] The DAPF module mainly includes:
[0037] First, the DAPF block shuffles the input x in the channel dimension and divides it into x 1:d and x d+1:D Two parts, and calculated according to the following formula to achieve distribution mapping:
[0038]
[0039] Where ⊙ denotes element-wise multiplication, s(·) and t(·) are separated by channel dimension from the output of the Deformable Attention based ST Transform Subnet (DA-STNet), and SCE is given by SCE(·) = exp(σ α (·)) is calculated, where:
[0040]
[0041] This formula means that before calculating exp(), we first pass σ α Add soft limits to the output to control the output within the range (-α,α).
[0042] Second, the output y of the first half 1:d and y d+1:D Perform the reverse calculation process, that is:
[0043]
[0044] Third, z 1:d With z d+1:D Merge in the channel dimension to obtain the final output z of the DAPF block.
[0045] In particular, the structure of DA-STNet can be expressed as:
[0046]
[0047] Among them, Conv2d represents a 3x3 two-dimensional convolution, ReLU represents a ReLU nonlinear activation function, LN represents layer normalization, and DA is the attention mechanism based on the deformable convolution proposed in the application, in which mechanism, first, a depthwise separable convolution (Depthwise Separable Convolution, DW Conv) block is used to preliminarily process the input feature x0 to obtain x1, so that the number of feature map channels is reduced to 1 / 4 of the original; then the efficient deformable convolution (Deformable Convolution) DCNv4 dynamically adjusts the convolution sampling position; then another DW Conv block is used to process to obtain x2, and the number of channels is restored to the same as that of x0; finally, the Sigmoid activation function is used to calculate the weighting coefficient of x2 and weight x0.
[0048] The multi-scale feature modeling stage uses a dynamic multi-scale feature fusion flow (DMFF) module to cooperatively model multi-scale features, and designs a dynamic multi-scale feature fusion based ST transform subnet (DMFF-STNet) for multi-scale transformation.
[0049] The input of the DMFF module is a multi-scale feature list, and the scaling parameter list s_list and the translation parameter list t_list in the intermediate process are directly obtained from the output of the dynamic multi-scale feature fusion based ST transform subnet DMFF-STNet
[0050] The DMFF-STNet first uses a multidimensional collaborative attention (MCA) module to respectively perform attention-based enhancement on the feature maps of each scale, to dynamically capture important information in different scale feature maps, and then performs a multi-scale feature fusion process.
[0051] The MCA module generates a weighting coefficient x_attn from the input x through multidimensional collaborative attention, and weights the feature map x_conv processed by the convolution block (composed of Conv2d, LN, and ReLU) (element-wise multiplication of x_attn and x_conv), and finally adjusts the weighted result through another Conv2d and outputs.
[0052] The multi-scale feature fusion process is as follows: first, the outputs of the MCA modules at the three scales are denoted as (The smaller j is, the shallower the feature map is, the higher the spatial resolution is and the fewer the number of channels is); then the feature map with the smallest spatial resolution is Adjust to the feature map of the previous layer through the upsampling convolution block (UpConv) The same size, denoted as
[0053] UpConv consists of a series of bilinear upsampling and Conv2d; then the dynamic feature fusion (DFF) module is used to fused the and Perform feature fusion to obtain Again Repeat the above The processing process is combined with Feature fusion Then output the three scales separately Each is processed by independent Conv_s and Conv_t (both Conv_s and Conv_t are 3×3 convolutions). Finally, the feature maps of the three scales obtained by Conv_s are combined (zip) into a list and recorded as s_list. The feature map list corresponding to Conv_t is recorded as t_list and used as the final output of DMFF-STNet.
[0054] The DFF module first performs Conv2d and LN operations on the module input x and y to obtain the intermediate feature map x mid and y mid , and multiply the two element by element to obtain the intermediate feature map z measured by the input feature similarity, and use the summation function σ to first sum z in the channel dimension, and then generate the attention weight of a single channel through the Sigmoid activation function, and finally perform feature fusion and output x in a similar way to the residual connection. out :
[0055] x out =x+σ(z)⊙y,
[0056] Furthermore, the overall optimization goal of the D3SFlow model in step S3 includes the following two parts:
[0057] First, the primary optimization goal is to train the model to map the normal features with unknown distribution to the multivariate standard Gaussian distribution in the latent space as much as possible, by maximizing the likelihood of the normal features in the feature space. First, when using the NF based on the affine coupling layer for distribution mapping, the log-likelihood of the input feature x can be preliminarily calculated by the following formula:
[0058] log p X (x) = log pZ (f(x))+∑ j s(x 1:d ) j ,
[0059] Among them, the transformation function f(x) represents the overall feature mapping process, which consists of two parts: single-scale modeling and multi-scale modeling. z is the output of the input feature x after being mapped by f(x), and p Z () is the probability likelihood of the mapping result z, ∑ j It represents the sum of the scaling parameters s output by DAPF-STNet and DMFF-STNet in single-scale modeling and multi-scale modeling.
[0060] Since the mapping result z∈R d It obeys the multivariate standard Gaussian distribution in the latent space, and its characteristic distribution expression is: Substituting it into the above formula, the constant part that does not affect the optimization objective is omitted. Therefore, rearrange the formula expression to obtain the loss function of the first part The initial expression:
[0061]
[0062] Among them, the subscript NLL indicates that the goal of this part of the loss is to minimize the input normal feature x N The sum of the negative log-likelihoods, N represents the total number of normal samples, D is the set of all normal features, for each negative log-likelihood -log p X (x N ), the first part of the formula represents the mapping result z N The square of the L2 norm can be obtained by z N The second part of the formula represents the sum of the scaling parameters of the ST conversion subnetwork output in the affine coupling layer contained in the NF, x N ' is a unified reference to the ST conversion subnet input.
[0063] According to the above description of the NF distribution mapping module in D3SFlow, the single-scale feature modeling stage uses an independent single-scale flow to model the feature map of each scale, and the output feature is represented as The number of DAPF modules included in each scale is The sum of the scaling parameters of the affine coupling layer output of the j-th single-scale stream is denoted as Among them are:
[0064]
[0065] Among them, s(x 1:d ) k and s(y d+1:D) k They represent the scaling parameters of the two DA-STNet outputs in the k-th DAPF module, respectively. j Indicates the number of DAPF modules used for the feature map of the jth scale. Similarly, the multi-scale feature modeling stage uses multi-scale flow DMFF to model the input of multiple scales simultaneously, and the output is expressed as The output scaling parameters of the two DMFF-STNets in the DMFF block are both multi-scale feature lists, denoted as s_list 1:d and s_list d+1:D , the sum of the scaling parameters S output by its affine coupling layer multi It can be calculated by the following formula:
[0066] S multi =∑ j [sum(s_list 1:d ) j +sum(s_list d+1:D ) j ],
[0067] Among them, sum(s_list) j It means that the j-th feature map in s_list is summed according to the sample dimension, and ∑j means that the sum of each item in s_list is added to obtain the sum of the scaling parameters of the affine coupling layer of this part S multi . Use the above expression to Rewriting this gives:
[0068]
[0069] Second, the secondary optimization goal is to optimize the mapping of pseudo-abnormal features so that they are as far away from the mapping results of normal features as possible, so as to alleviate the problem of low distinguishability of the mapping results of abnormal features that are invisible during the training process. It is possible to obtain each normal feature x in the sample dimension N Negative log-likelihood -logp X (x N ), extended to pseudo-abnormal features, each pseudo-abnormal feature x can be obtained in the same way A The negative log-likelihood of -logpX(x A ). In order to guide the model to focus on the optimization of normal features as the main optimization goal, and adjust the mapping of pseudo-abnormal features as a secondary goal, this part of the loss The optimization goal is to increase the negative log-likelihood of the pseudo-abnormal feature until it is greater than the optimization boundary B, that is, to keep a certain distance from the negative log-likelihood of the normal feature so that the model can distinguish between the two. Specifically, first calculate the maximum negative log-likelihood of the normal samples in the batch, denoted as -log p max , and then set the optimization boundary B of the pseudo-anomaly samples in the batch to:
[0070]
[0071] Among them, τ is the minimum distance between the negative log-likelihood of the preset normal features and pseudo-abnormal features. Correspondingly, the second part of the model optimization objective It can be calculated as follows:
[0072]
[0073] The subscript CLO indicates that the goal of this part of the loss is to optimize the collaborative negative log-likelihood of normal features and pseudo-abnormal features until the distance between the two is greater than τ, N represents the total number of pseudo-abnormal samples, and A is the set of all pseudo-abnormal features. The total loss function is then obtained by the following formula: (λ is the weight of the two losses added together):
[0074]
[0075] Furthermore, the calculation of the anomaly score in step S4 includes the following two calculation steps:
[0076] (1) Pixel-level anomaly score calculation: First, calculate the log-likelihood of the mapping results at each scale, take the average according to the channel dimension, and then upsample it to the input image resolution in terms of spatial resolution. Then, process the upsampled results and use the likelihood in the range of (0,1). Perform multi-scale aggregation and use the accumulation and multiplication method to obtain;
[0077]
[0078] Finally, the pixel-level anomaly score map S is generated based on the aggregation results in the following way. loc :
[0079] S loc =S add =max(P add )-P add ,
[0080] (2) Image-level anomaly score calculation. The specific calculation formula is as follows:
[0081]
[0082] First, multiply the result P by multi-scale likelihood mul Calculate the anomaly score map S mul , and then through the abnormal score map S mul The average of the first K pixels with the largest anomaly scores in the image is the final image-level anomaly score s det .
[0083] Furthermore, the present invention also proposes an industrial image anomaly detection system based on normalized flow, which applies the above-mentioned industrial image anomaly detection based on flow model and includes two main functional modules:
[0084] (1) Visual dataset analysis module: Based on the dataset selected by the user, the module can intuitively display the statistical information of the dataset in the form of visual charts, including the number of positive and abnormal images of different products in the dataset. For each product, the module can also display the statistical information of the average size of a single product and the sizes of different abnormal areas.
[0085] (2) Model online reasoning and result display module: It can match and load the offline model of the corresponding product type into the system according to the user's settings (including data set, product type and specific image to be tested), and use the D3SFlow model designed by the present invention to complete the online reasoning of the image to be tested, and save the image of the detection result, and finally visually display the model reasoning result on the user page.
[0086] Compared with the prior art, the present invention has the following beneficial effects:
[0087] (1) The present invention introduces the self-supervised paradigm as an optimization method and adds the process of explicitly learning the distribution of pseudo-abnormal features to the model, aiming to improve the problem of low distinguishability of the positive abnormal feature mapping results generated by NF in the model during feature mapping. It proposes a pseudo-abnormal generation strategy in the image dimension and a pseudo-abnormal generation strategy in the feature dimension, realizes the generation of pseudo-abnormal data from two dimensions, and designs a self-supervised loss function based on collaborative likelihood optimization as the loss function of the model, so that the model can simultaneously learn how to map normal features to multivariate Gaussian distribution and how to map the mapping results of pseudo-abnormal features that deviate from normal features, which can fit the normal feature distribution more accurately and improve the model's ability to distinguish positive and abnormal features.
[0088] (2) The application proposes a flow model composed of two stages for feature modeling. In the first stage, a parallel flow model based on deformable convolution attention (DAPF) is used to learn feature mapping in a single scale for different scale features respectively. In the second stage, a flow model based on dynamic multi-scale feature fusion (DMFF) is used to learn feature mapping in multiple scales simultaneously. This scheme not only ensures the multi-scale modeling capability of the model, but also reduces the computational amount of the model. In addition, the application also adopts a feature preprocessing strategy based on learnable downsampling, which performs downsampling in the feature preprocessing stage after extracting pre-trained features and before inputting the NF learning feature mapping, further reducing the parameter amount of the model while trying to retain the anomaly detection capability of the model.
[0089] (3) Extensive comparative experiments were conducted on three public datasets MVTec-AD, BTAD and VisA. The D3SFlow model proposed in the application was compared with single-class industrial image anomaly detection models based on different principles in the past three years, and competitive image-level anomaly detection (I-AUROC) and pixel-level anomaly localization (P-AUROC and AUPRO) results were obtained. Through the visualization of the detection results, the applicability and accurate anomaly localization capability of the D3SFlow model in various product anomaly detection tasks were verified.
[0090] (4) The algorithm prototype system developed based on the model of the application provides dataset visualization analysis and model online reasoning functions. The dataset visualization analysis function can intuitively display the statistical information of the dataset, helping users better understand the data characteristics. The model online reasoning function is simple to operate and the result display is clear, which can meet the actual use needs of users in industrial image anomaly detection and has good application prospect, providing strong support for quality control in industrial production. BRIEF DESCRIPTION OF DRAWINGS
[0091] Figure 1 Flowchart of the industrial image anomaly detection method based on the flow model.
[0092] Figure 2 Structure diagram of the feature extraction module.
[0093] Figure 3 Structure diagram of the parallel DAPF module based on deformable convolution attention.
[0094] Figure 4 Structure diagram of the interactive flow DMFF module based on dynamic multi-scale feature fusion.
[0095] Figure 5 Structure diagram of the ST conversion subnetwork DMFF-STNet based on dynamic multi-scale feature fusion. DETAILED DESCRIPTION
[0096] The application will be further described below with reference to the accompanying drawings.
[0097] The industrial image anomaly detection method based on a self-supervised flow model has the overall flow as shown in FIG. 1, and the specific steps include: Figure 1
[0098] S1. The training image is sequentially input into the feature extraction module and the feature preprocessing module (the resolution of the training image is uniformly set to 512*512 in this embodiment), and the feature maps (multi-scale normal anomaly feature maps) containing normal features and pseudo-anomaly features at three different scales are output.
[0099] The overall structure of the feature extraction module is shown in FIG. 2, and according to the difference in the dimension of the pseudo-anomaly generation strategy used, the structure of one of the upper part or the lower part can be selected in the training stage. The upper part is the module structure under the image dimension, the normal and pseudo-anomaly images are obtained through image dimension pseudo-anomaly generation (I-PAG), and then pre-training feature extraction is performed to obtain multi-scale normal anomaly features; the lower part is the module structure under the feature dimension, the multi-scale normal features are obtained through pre-training feature extraction, and then pseudo-anomaly features are obtained through feature dimension pseudo-anomaly generation (F-PAG). In the inference stage, no pseudo-anomaly generation strategy is used, and only the pre-training feature extraction part common to both is retained. Figure 2 The image dimension pseudo-anomaly generation process can be divided into the following steps:
[0100] (1) Generate a pseudo-anomaly region mask through the foreground-enhanced Berlin noise mask generator;
[0101] (2) Generate an image as an anomaly source through the structure and texture anomaly generator, generate a pseudo-structure anomaly image of a homologous data set with a probability p, and generate a pseudo-texture anomaly image from a heterologous data set with a probability of 1-p (in this embodiment, p is 0.5);
[0102] (3) First, the normal region mask is generated by reversing the pixel value of the anomaly region mask generated by the foreground-enhanced Berlin noise mask generator, and then the final simulated anomaly image is generated using the following formula:
[0103]
[0104]
[0105] wherein M is the pseudo-anomaly region mask generated by the foreground-enhanced Berlin noise mask generator, M is a normal area mask generated by inverting pixel values, N is a normal image to be generated as a pseudo-anomaly, A is an abnormal source image generated by the structure and texture anomaly generator, and β is an abnormality degree factor with a value interval of (0.2, 0.8). In this embodiment, the value of β is generated with random probability within the above interval, indicating that an abnormal source is added to the abnormal area at a ratio of β, and the original image is retained at a ratio of 1-β.
[0106] The feature dimension pseudo-anomaly generation process can be divided into the following steps:
[0107] (1) Generate abnormal area masks (M1, M2, M3) corresponding to feature maps of different sizes through the pyramid mask generator;
[0108] (2) Through the feature perturbation of foreground enhancement, the corresponding masks (M1, M2, M3) are used to add noise perturbation to the feature maps (h1, h2, h3) of different scales according to the following formula to obtain the corresponding pseudo-abnormal feature h A , where ∈ is independently sampled from a Gaussian distribution
[0109] h A =h N +M i ⊙∈
[0110] Among them, h N It is one of the normal feature maps (h1, h2, h3) processed by the feature perturbation of foreground enhancement, M i is the abnormal area mask (M1, M2, M3) generated by the pyramid mask generator and h N Corresponding masks of the same size;
[0111] (3) Concatenate normal features and pseudo-abnormal features in the batch dimension and output a multi-scale feature map containing both normal features and pseudo-abnormal features (where the batch size is twice the size of the input feature map).
[0112] In the selection of pre-trained networks, the model needs to have comprehensive detection capabilities for products of different sizes and anomalies of different scales. Considering that the feature maps of the pre-trained ViT network consider more global correlations and lack sufficient attention to small-scale anomalies in the dataset, which is crucial in industrial image anomaly detection, the pre-trained network of the CNN structure naturally meets the following characteristics: the feature maps of different stages have different receptive fields (the more convolutional layers, the larger the receptive field of the feature map). By combining the feature maps of different stages, the above detection requirements can be well met. In this embodiment, WideResNet-50 is used as the CNN pre-training network Φ for pre-training feature extraction.
[0113] The steps of pre-training feature extraction are as follows: the dataset is divided into training set X according to the requirements of the single-category industrial anomaly detection (OC-IAD) task. train and the test set X test , X train ∪X test The image x in i Input the fixed parameter CNN pre-trained network φ and sample the last feature map of multiple stages of the network Where L is set to 3 to extract the feature maps of the first three stages of φ, j is the index of the feature maps of different stages, and H j 、W j 、C j Represent the height, width and channel dimension of the feature map of the jth stage respectively, so the pre-training feature extraction process can be expressed as:
[0114] (h1,h2,h3)=Φ(x)
[0115] Here, φ(·) represents a pre-trained network that follows the multi-scale feature extraction rules described above. In this embodiment, the feature maps h1, h2, and h3 from the first three stages of the pre-trained network are selected as feature maps of different scales for subsequent modeling because they contain more spatial detail information, which is more important for detecting small-scale anomalies. The feature map output from the last stage has too low a resolution and can provide limited assistance for anomaly localization. It should also be noted that the parameters of the pre-trained network are always frozen during training.
[0116] The feature preprocessing module includes the following steps:
[0117] (1) The input feature map is downsampled by the learnable downsampling (LDS) module. Specifically, the input feature is first decomposed into the following components: the diagonal high-frequency component y D , high-frequency component y in the vertical direction V , high-frequency component y in the horizontal direction H , low-frequency component y A , and concatenate them in the channel dimension, and then align the number of channels through the learnable feature alignment (LFA) module;
[0118] (2) Since the input industrial product images and the natural images used in pre-training come from different domains, there are differences in the feature distribution characteristics between the two, which may affect the effectiveness of the pre-trained features in the industrial anomaly detection task. Therefore, a feature adapter (FA) module is used to map the pre-trained features h' after downsampling by the LDS module to the target domain, where FA is composed of a single fully connected layer.
[0119] S2. Input the multi-scale positive anomaly feature map obtained in S1 into the Normalized Flow (NF) distribution mapping module, mapping the input features in the original feature space (distribution unknown) to the multivariate standard Gaussian distribution in the latent space (distribution known), so that the model can distinguish positive anomaly features based on the likelihood value of the input features after being mapped by this module;
[0120] The NF distribution mapping module uses a carefully designed two-stage flow model. The first stage is the single-scale feature modeling stage, which is mainly responsible for learning the intrinsic properties of a single-scale feature map, that is, independently learning each scale feature map through three parallel NFs to maximize the mapping method of likelihood. The second stage is the multi-scale feature modeling stage, which is mainly responsible for sharing the global semantic information and local detail information contained in feature maps of different scales, that is, learning a mapping method that maximizes the comprehensive likelihood of the input feature map at multiple scales through a single NF, so that the model can simultaneously utilize different effective information in multi-scale feature maps for anomaly detection.
[0121] During the single-scale feature modeling phase of the stream model, the Deformable Convolutional Attention Parallel Flow (DAPF) module is used to independently model feature maps at three scales. To improve distribution mapping efficiency and reduce the overall computational complexity of the model, a pyramid-like structure is adopted, using different numbers of DAPF modules (proportional to the number of channels of the feature) to distribute and map features at different scales. Specifically, since the three feature maps (h1, h2, h3) at this stage have 256, 512, and 1024 channels, respectively, a corresponding number of DAPF modules (1, 2, and 4) are used for distribution mapping.
[0122] The DAPF module is composed of two consecutive affine coupling layers of opposite transformation processes. Figure 3 The upper part shows the following:
[0123] First, the DAPF module first merges the original input single-scale feature map x and the position code containing absolute position information generated by the position encoder (PE) (the number of channels of the position code output is uniformly set to 64 in this embodiment) in the channel dimension as the overall input of the DAPF module, and then shuffles the input in the channel dimension and divides it into x 1:d and x d+1:D Two parts, and calculated according to the following formula to achieve distribution mapping:
[0124]
[0125] where, represents element-wise multiplication, subscript D represents the total number of channels of the feature map, s(·) and t(·) are separated from the output of the ST transform subnetwork based on the deformable convolution attention (DA-STNet) in the channel dimension, and SCE is calculated by SCE(·) = exp(σ α (·)) where there is:
[0126]
[0127] The above formula indicates that σ α A soft limit is added to the output to control the output in the range of (-α, α), which increases the training stability, and α is 3 in the present model.
[0128] Second, the output y 1:d and y d+1:D of the first half are calculated in the opposite way to ensure that all channels can pass through the mapping operation, which is expressed as:
[0129]
[0130] Third, z 1:d and z d+1:D are merged in the channel dimension to obtain the final output z of the DAPF module.
[0131] where the structure of the DA-STNet is shown in the lower left part of the attached Figure 3 and can be expressed as:
[0132]
[0133] where Conv2d represents a two-dimensional 3x3 convolution (to preserve the spatial information of the input feature map), the number of intermediate feature channels of the two Conv2ds is the same as the input, and the number of output feature channels is twice the input, because the output of the DA-STNet needs to be separated into two feature maps s(x) and t(x) with the same size as the input feature, ReLU represents the ReLU nonlinear activation function, and DA is the deformable convolution-based attention mechanism proposed in the present application, and the specific structure is shown in the attached Figure 3Specifically, the input feature x0 is first processed using a depthwise separable convolution (DW Conv) block (followed by a series LN and ReLU to improve stability and enhance nonlinear representation capabilities) to obtain x1, reducing the number of feature map channels to 1 / 4 to reduce computational complexity. The highly efficient deformable convolution (DCNv4) is then used to dynamically adjust the convolution sampling position, enabling the model to adaptively focus on more important areas, thereby enhancing the model's ability to model complex feature maps. Another DW Conv block is then used to obtain x2, restoring its number of channels to the same as x0 to maintain the shape of the feature map. Finally, a sigmoid activation function is used to calculate the weighting coefficient for x2 and weight x0 to achieve selective attention to key areas.
[0134] In the multi-scale feature modeling stage of the flow model, a dynamic multi-scale feature fusion flow (DMFF) module is used for multi-scale feature collaborative modeling, and an ST transformation subnet based on dynamic multi-scale feature fusion (DMFF-STNet) is designed for multi-scale transformation.
[0135] The structure of the DMFF module is shown in the attached Figure 4 As shown in the figure, the DMFF block, similar to the DAPF module used in the single-scale feature modeling stage, also adopts the NF structure based on the affine coupling block. However, its input is a multi-scale feature list, and the intermediate scaling parameter list s_list and translation parameter list t_list are directly derived from the output of DMFF-STNet. DMFF-STNet can dynamically fuse the input multi-scale feature maps, giving the model the ability to flexibly utilize the different information contained in the multi-scale feature maps for anomaly detection.
[0136] The specific structure of DMFF-STNet is shown in the attached Figure 5 As shown in the left half of the figure, the multi-dimensional collaborative attention (MCA) module is first used to perform attention-based enhancement on the feature maps of each scale to dynamically capture important information in the feature maps of different scales. The specific structure of the MCA module is shown in the attached Figure 5 As shown in the lower right part of the figure, the input x is subjected to multi-dimensional collaborative attention to generate weighted coefficients x_attn. The feature map x_conv processed by the convolutional block (consisting of Conv2d, LN, and ReLU) is weighted (by multiplying x_attn and x_conv element-wise). Finally, another Conv2d performs the final adjustment on the weighted result and outputs it. This is followed by the multi-scale feature fusion process.
[0137] The process of multi-scale feature fusion is as follows: First, the output of the MCA module at three scales is recorded as (The smaller j is, the shallower the feature map is, the higher the spatial resolution is and the fewer the number of channels is); then the feature map with the smallest spatial resolution is Adjust to the feature map of the previous layer through the upsampling convolution block (UpConv) The same size, denoted as UpConv consists of a series of bilinear upsampling and Conv2d; then the dynamic feature fusion (DFF) module and Perform feature fusion to obtain Again Repeat the above The processing process is combined with Feature fusion Then output the three scales separately Each is processed by independent Conv_s and Conv_t (both Conv_s and Conv_t are 3×3 convolutions). Finally, the feature maps of the three scales obtained by Conv_s are combined (zip) into a list and recorded as s_list. The feature map list corresponding to Conv_t is recorded as t_list and used as the final output of DMFF-STNet.
[0138] The DFF module can flexibly fuse the features of two input feature maps of different scales, so that the model can better utilize information from different scales for anomaly detection. Its specific structure is shown in the attached figure. Figure 5 As shown in the upper right part of the figure, the DFF module first processes the module input x and y through Conv2d and LN operations respectively to obtain the intermediate feature map x mid and y mid , and multiply the two element by element to obtain the intermediate feature map z measured by the input feature similarity, and use the summation function σ to first sum z in the channel dimension, and then generate the attention weight of a single channel through the Sigmoid activation function, and finally perform feature fusion and output x in a similar way to the residual connection. out , the above process is expressed using the following formula:
[0139] x out =x+σ(z)⊙y,
[0140] The aforementioned DMFF-STNet selectively and dynamically fuses the input multi-scale feature maps, enabling the model to comprehensively learn the information of different granularities contained in the multi-scale feature maps without interfering with each other. Therefore, the DMFF, of which it is a key component, can fully complete multi-scale feature modeling. The NF distribution mapping module, as a whole, adopts a strategy of first modeling single-scale features and then multi-scale models. This ensures that, after learning the unique information of a single scale, the model uses the shared information of multiple scales to strengthen its modeling capabilities for the input multi-scale features, thereby improving the overall anomaly detection performance of D3SFlow.
[0141] S3. Update model parameters by minimizing the self-supervised loss function based on collaborative likelihood optimization (CLO-SSL), optimize the distribution of normal features and pseudo-anomaly features, and gradually approach the overall optimization goal of the D3SFlow model;
[0142] The overall optimization objectives of the model are divided into the following two points:
[0143] First, the primary optimization goal is to train the model to map the normal features with unknown distribution to the multivariate standard Gaussian distribution in the latent space as much as possible, by maximizing the likelihood of the normal features in the feature space. First, when using the NF based on the affine coupling layer for distribution mapping, the log-likelihood of the input feature x can be preliminarily calculated by the following formula:
[0144] log p X (x) = log p Z (f(x))+∑ j s(x 1:d ) j ,
[0145] Among them, the transformation function f(x) represents the overall feature mapping process, which consists of two parts: single-scale modeling and multi-scale modeling. z is the output of the input feature x after being mapped by f(x), and p Z () is the probability likelihood of the mapping result z, Σ j It represents the sum of the scaling parameters s output by DAPF-STNet and DMFF-STNet in single-scale modeling and multi-scale modeling.
[0146] Since the mapping result z∈R d It obeys the multivariate standard Gaussian distribution in the latent space, and its characteristic distribution expression is: Substituting it into the above formula, the constant part that does not affect the optimization objective is omitted. Therefore, by rearranging the formula expression, the loss function of the first part can be obtained The initial expression:
[0147]
[0148] Among them, the subscript NLL indicates that the goal of this part of the loss is to minimize the input normal feature x N The sum of the negative log-likelihoods, N represents the total number of normal samples, D is the set of all normal features, for each negative log-likelihood -log p X (x N ), the first part of the formula represents the mapping result z N The square of the L2 norm can be obtained by z N The second part of the formula represents the sum of the scaling parameters of the ST conversion subnetwork output in the affine coupling layer contained in the NF, x N ' is a unified reference to the ST conversion subnet input.
[0149] According to the above description of the NF distribution mapping module in D3SFlow, the single-scale feature modeling stage uses an independent single-scale flow to model the feature map of each scale, and the output feature is represented as The number of DAPF modules included in each scale is The sum of the scaling parameters of the affine coupling layer output of the j-th single-scale stream is denoted as Among them are:
[0150]
[0151] Among them, s(x 1:d ) k and s(y d+1:D ) k They represent the scaling parameters of the two DA-STNet outputs in the k-th DAPF module, respectively. j Indicates the number of DAPF modules used for the feature map of the jth scale. Similarly, the multi-scale feature modeling stage uses multi-scale flow DMFF to model the input of multiple scales simultaneously, and the output is expressed as The output scaling parameters of the two DMFF-STNets in the DMFF block are both multi-scale feature lists, denoted as s_list 1:d and s_list d+1:D , the sum of the scaling parameters S output by its affine coupling layer multi It can be calculated by the following formula:
[0152] S multi =∑ j [sum(s_list 1:d ) j +sum(s_list d+1:D ) j ],
[0153] Among them, sum(s_list) jIt means that the j-th feature map in s_list is summed according to the sample dimension, and ∑j means that the sum of each item in s_list is added to obtain the sum of the scaling parameters of the affine coupling layer of this part S multi . Use the above expression to Rewriting this gives:
[0154]
[0155] Second, the secondary optimization goal is to train the model to map pseudo-abnormal features outside the preset distribution of the latent space, so as to further optimize the mapping of normal features and alleviate the problem of low distinguishability of the model's mapping results for invisible abnormal features. It is possible to obtain each normal feature x in the sample dimension N Negative log-likelihood -logp X (x N ), extended to pseudo-abnormal features, each pseudo-abnormal feature x can be obtained in the same way A The negative log-likelihood of -logpX(x A ). In order to guide the model to focus on the optimization of normal features as the main optimization goal, and adjust the mapping of pseudo-abnormal features as a secondary goal, this part of the loss The optimization goal is to increase the negative log-likelihood of the pseudo-abnormal feature until it is greater than the optimization boundary B, that is, to keep a certain distance from the negative log-likelihood of the normal feature so that the model can distinguish between the two. Specifically, first calculate the maximum negative log-likelihood of the normal samples in the batch, denoted as -log p max , and then set the optimization boundary B of the pseudo-anomaly samples in the batch to:
[0156]
[0157] Where τ is the minimum distance between the negative log-likelihood of the preset normal feature and the pseudo-abnormal feature, and in this embodiment, the value is 1e -5 Correspondingly, the second part of the model optimization objective It can be calculated as follows:
[0158]
[0159] The subscript CLO indicates that the goal of this part of the loss is to optimize the collaborative negative log-likelihood of normal features and pseudo-abnormal features until the distance between the two is greater than τ, N represents the total number of pseudo-abnormal samples, and A is the set of all pseudo-abnormal features. The total loss function is then obtained by the following formula: (λ is the weight of the two losses added together, and in this embodiment, the value is 0.5):
[0160]
[0161] S4. Input the image to be detected into the trained D3SFlow model, calculate the anomaly score of the model output, and generate image-level and pixel-level anomaly detection results.
[0162] The anomaly score calculation includes the following two calculation steps:
[0163] (1) Pixel-level anomaly score calculation: First, calculate the log-likelihood of the mapping results at each scale, take the average according to the channel dimension, and then upsample it to the input image resolution in terms of spatial resolution. Then, process the upsampled results and use the likelihood in the range of (0,1). Perform multi-scale aggregation and use the accumulation and multiplication method to obtain;
[0164]
[0165] Finally, the pixel-level anomaly score map S is generated based on the aggregation results in the following way. loc :
[0166] S loc =S add =max(P add )-P add ,
[0167] (2) Image-level anomaly score calculation. The specific calculation formula is as follows:
[0168]
[0169] First, multiply the result P by multi-scale likelihood mul Calculate the anomaly score map S mul , to prevent the accidental abnormal signals in the single scale mapping results from affecting the overall image-level abnormality detection results, and then through the abnormal score map S mul The average of the first K pixels with the largest anomaly scores in the image is the final image-level anomaly score s det ,In this embodiment, the value of K is 3% of the total number of pixels in the test image, ,in order to take into account both the sensitivity of abnormal pixel detection and the ,stability of the overall image-level indicators.
[0170] In terms of model training parameter settings, this embodiment adopts the following settings: the batch size is set to 8, the Adam optimizer is used for training, the number of training rounds (epochs) is 100, the learning rate adopts the warmup strategy, and the initial learning rate is set to 1e -5 , and gradually increase to the base learning rate 1e within 5 rounds -4 .
[0171] This embodiment uses the D3SFlow under the above settings for the single-class industrial anomaly detection (OC-IAD) task on the MVTec-AD, BTAD, and VisA datasets, that is, training a model for each product category. The evaluation indicators used in this embodiment are AUROC (Area Under the Receiver Operating Curve) and AUPRO (Area Under the PerRegion Overlap). AUROC measures the performance of the model using the area under the ROC curve (a curve drawn by the true positive rate (TPR) and the false positive rate (FPR) at different thresholds), the higher the value, the less the impact of different thresholds on the performance of the model, and the stronger the ability of the model to distinguish between normal and abnormal categories. This embodiment measures from both image-level AUROC (I-AUROC) and pixel-level AUROC (P-AUROC); AUPRO is similar to the IoU (Intersection Over Union) indicator commonly used in semantic segmentation, which uses the area under the PRO curve (calculated by the overlap of each connected region with the ground truth mask) and calculates the average pixel false positive rate of 30% on the entire dataset to measure the abnormal positioning performance of the model.
[0172] This embodiment uses the above evaluation indicators to conduct experiments on 15 product categories (10 object type products and 5 texture type products) in the MVTec-AD dataset, achieving an average performance of 99.8% / 98.8% / 97.9% on the I-AUROC / P-AUROC / AUPRO indicators (calculated after evaluating each product category separately); for the 3 complex texture product categories in the BTAD dataset, the average performance on the I-AUROC / P-AUROC / AUPRO indicators is 97.2% / 98.1% / 87.2%; for the 4 single-instance products, 4 multi-instance products, and 4 complex structure products in the VisA dataset, the average performance on the I-AUROC / P-AUROC / AUPRO indicators is 98.5% / 99.1% / 96.0%. The experiments on the above three datasets show that the D3SFlow model has high detection accuracy on the OC-IAD task, whether it is image-level anomaly detection ability or pixel-level anomaly positioning.
Claims
1. A method for detecting anomalies in industrial images based on a self-supervised flow model, characterized in that: This involves building a self-supervised flow model with dynamic feature selection (D3SFlow), which includes a feature extraction module, a feature preprocessing module, and a normalized flow (NF) distribution mapping module. This is followed by model training and industrial image anomaly detection. The specific steps are as follows: S1. Input the training image into the feature extraction module and the feature preprocessing module in sequence to obtain feature maps at three different scales containing normal features and pseudo-abnormal features, namely, multi-scale positive and abnormal feature maps; where: The feature extraction module has two different structures according to the dimensions of the pseudo-anomaly generation strategy used: one is a module structure under the image dimension, which first obtains normal and pseudo-anomaly images through the image dimension pseudo-anomaly generation step, and then performs pre-training feature extraction to obtain multi-scale positive anomaly features; the other is a module structure under the feature dimension, which first obtains multi-scale normal features through pre-training feature extraction, and then obtains pseudo-anomaly features through the feature dimension pseudo-anomaly generation step; The feature preprocessing module has the following workflow: (1) The input feature map is downsampled through a learnable downsampling module. Specifically, the input feature is first decomposed into the following components: the diagonal high-frequency component y D , high-frequency component y in the vertical direction V , high-frequency component y in the horizontal direction H , low-frequency component y A , and splice them in the channel dimension, and then align the number of channels through the learnable feature alignment module; (2) Use the feature adapter module to map the downsampled pre-trained features to the target domain; S2. Input the multi-scale positive and abnormal feature map obtained in S1 into the NF distribution mapping module, map the input features in the original feature space (distribution unknown) to the multivariate standard Gaussian distribution in the latent space (distribution known), and use the log-likelihood of the input features after mapping by this module to distinguish positive and abnormal features; The workflow of the NF distribution mapping module is as follows: (1) Single-scale feature modeling, which is responsible for learning the intrinsic properties of a single-scale feature map, that is, independently learning the mapping method that maximizes the likelihood of each scale feature map through three parallel NF distributions; (2) Multi-scale feature modeling, which is responsible for sharing the global semantic information and local detail information contained in feature maps of different scales. That is, a mapping method that maximizes the comprehensive likelihood of the input feature map at multiple scales is learned through a single NF, so that the model can use different effective information in the multi-scale feature map for anomaly detection; S3. Update model parameters by minimizing the self-supervised loss function based on collaborative likelihood optimization, optimize the distribution of normal features and pseudo-anomaly features, and gradually approach the overall optimization goal of the D3SFlow model. The optimization goal consists of the following two parts: (1) The primary optimization goal is to train the model to map normal features with unknown distribution to the multivariate standard Gaussian distribution in the latent space as much as possible, which is achieved by maximizing the log-likelihood of normal features in the feature space; (2) The secondary optimization goal is to optimize the mapping of pseudo-anomaly features so that they are as far away from the mapping results of normal features as possible, so as to alleviate the problem of low distinguishability of the mapping results of abnormal features that are invisible during the training process; S4. Input the image to be detected into the trained D3SFlow model, calculate the anomaly score of the model output, and generate image-level and pixel-level anomaly detection results; Among them, S1-S3 are the training stage, and S4 is the detection stage.
2. The industrial image anomaly detection method according to claim 1, characterized in that: In step S1: The image dimension pseudo-anomaly generation process is divided into the following steps: (1) Generate pseudo-abnormal region masks through foreground-enhanced Perlin noise mask generator; (2) Generate images as anomaly sources through structural and texture anomaly generators, generate pseudo structural anomaly images from the same source dataset with probability p, and generate pseudo texture anomaly images from the heterogeneous dataset with probability 1-p; (3) First, the abnormal area mask generated by the foreground enhanced Perlin noise mask generator is converted into a normal area mask by inverting the pixel values, and then the final simulated abnormal image is generated using the following formula: Among them, M is the pseudo abnormal area mask generated by the foreground enhanced Perlin noise mask generator, M is the normal area mask generated by pixel value inversion, N is the normal image to be generated as a pseudo-anomaly, A is the abnormal source image generated by the structure and texture anomaly generator, β is the abnormality degree factor with a value range of (0.2, 0.8), which means that the abnormal source is added to the abnormal area at a ratio of β and the original image is retained at a ratio of 1-β; The feature dimension pseudo-anomaly generation process is divided into the following steps: (1) Generate abnormal area masks (M1, M2, M3) corresponding to feature maps of different sizes through the pyramid mask generator; (2) Through the feature perturbation of foreground enhancement, the corresponding masks (M1, M2, M3) are used to add noise perturbation to the feature maps (h1, h2, h3) of different scales according to the following formula to obtain the corresponding pseudo-abnormal feature h A : h A =h N +M i ⊙∈, Among them, h N It is one of the normal feature maps (h1, h2, h3) processed by the feature perturbation of foreground enhancement, M i is the abnormal area mask (M1, M2, M3) generated by the pyramid mask generator and h N Corresponding masks of the same size; ∈ independently sampled from Gaussian distribution (3) Concatenate normal features and pseudo-anomaly features in the batch dimension and output a multi-scale feature map containing both normal features and pseudo-anomaly features. The batch size output by this module is twice the size of the input feature map. In the pre-training feature extraction process, the dataset is divided into a training set and a test set according to the requirements of the single-category industrial anomaly detection (OC-IAD) task. The images in the union of the two sets are input into a CNN pre-trained network Φ with fixed parameters, and the last feature map of multiple stages of the network is sampled. Where L=3 means extracting the feature maps of the first three stages, so the pre-training feature extraction process is expressed as the following formula: (h1,h2,h3)=Φ(x), Among them, h1, h2, and h3 represent feature maps of three different scales respectively; here Φ adopts WideResNet-50.
3. The industrial image anomaly detection method according to claim 2, characterized in that: In step S2: In the single-scale feature modeling stage, the Deformable Convolution Attention Parallel Flow (DAPF) module is used to independently model the feature maps of the three scales. In order to improve the efficiency of distribution mapping and reduce the overall computational complexity of the model, a pyramid-like structure is used to distribute and map features of different scales using different numbers of DAPF modules. Specifically: First, the input x is shuffled in the channel dimension and divided equally into x 1:d and x d+1:D The two parts are calculated according to the following formula to achieve distribution mapping: Where ⊙ represents element-wise multiplication, s(·) and t(·) are separated by channel dimension from the output of the deformable convolutional attention-based ST transformation subnetwork (DA-STNet), and SCE is given by SCE(·) = exp(σ α (·)) is calculated, where: This formula means that before calculating exp(), we first pass σ α Add soft limits to the output to control the output within the range of (-α,α); Second, the output y of the first half 1:d and y d+1:D Perform the reverse calculation process, that is: Third, z 1:d With z d+1:D Merge in the channel dimension to obtain the final output z of the DAPF module; The structure of the DA-STNet is represented as: Among them, Conv2d represents a two-dimensional 3×3 convolution, ReLU represents the ReLU nonlinear activation function, LN represents layer normalization, and DA is an attention mechanism based on deformable convolution. In this mechanism, the depth-wise separable convolution (DW Conv) block is first used to perform preliminary processing on the input feature x0 to obtain x1, reducing the number of feature map channels to 1 / 4 of the original; then the convolution sampling position is dynamically adjusted through the efficient deformable convolution DCNv4; then another DW Conv block is used to process x2, restoring its number of channels to the same as x0; finally, the weighting coefficient of x2 is calculated using the Sigmoid activation function and weighted with x0; In the multi-scale feature modeling stage, a dynamic multi-scale feature fusion flow (DMFF) module is used to perform multi-scale feature collaborative modeling, and an ST conversion subnet based on dynamic multi-scale feature fusion (DMFF-STNet) is designed to perform multi-scale transformation; The input of the DMFF module is a multi-scale feature list, and the scaling parameter list s_list and translation parameter list t_list of the intermediate process are directly obtained from the output of DMFF-STNet; The DMFF-STNet first uses the multi-dimensional collaborative attention (MCA) module to perform attention-based enhancement on the feature maps of each scale to dynamically capture important information in feature maps of different scales, and then performs a multi-scale feature fusion process; The MCA module generates a weighted coefficient x_attn for the input x through multi-dimensional collaborative attention, and weights the feature map x_conv processed by the convolution block composed of Conv2d, LN, and ReLU, that is, multiplying x_attn and x_conv element by element, and finally another Conv2d makes the final adjustment on the weighted result and outputs it.
4. The industrial image anomaly detection method according to claim 3, characterized in that: In step S2: The process of multi-scale feature fusion is as follows: First, the output of the MCA module at three scales is recorded as Then the feature map with the smallest spatial resolution Adjust to the feature map of the previous layer through the upsampling convolution block (UpConv) The same size, denoted as UpConv consists of a series of bilinear upsampling and Conv2d; then the dynamic feature fusion (DFF) module and Perform feature fusion to obtain Again Repeat the above The processing process is combined with Feature fusion Then output the three scales separately Each is processed by independent Conv_s and Conv_t. Both Conv_s and Conv_t are 3×3 convolutions. Finally, the feature maps of the three scales obtained by Conv_s are combined into a list and recorded as s_list. The feature map list corresponding to Conv_t is recorded as t_list and used as the final output of DMFF-STNet. The DFF module first performs Conv2d and LN operations on the input x and y respectively to obtain the intermediate feature map x mid and y mid And multiply the two element by element to obtain the intermediate feature map z measured by the input feature similarity, and use the summation function σ to sum z in the channel dimension, and then use the Sigmoid activation function to generate the attention weight of the single channel, and finally perform feature fusion through the residual connection and output x out : x out =x+σ(z)⊙y。 5. The industrial image anomaly detection method according to claim 4, characterized in that: In step S3: For the optimization objective (1), first, when using the NF based on the affine coupling layer for distribution mapping, the log-likelihood of the input feature x is preliminarily calculated by the following formula: log p X (x)=log p Z (f(x))+∑ j s(x 1:d ) j , Among them, the transformation function f(x) represents the overall feature mapping process, which consists of two parts: single-scale modeling and multi-scale modeling. z is the output of the input feature x after being mapped by f(x), and p Z () is the probability likelihood of the mapping result z, ∑ j Indicates the summation of the scaling parameters s output by DAPF-STNet and DMFF-STNet in single-scale modeling and multi-scale modeling; Since the mapping result z∈R d It obeys the multivariate standard Gaussian distribution in the latent space, and its characteristic distribution expression is: Substituting it into the above formula, omitting the constant part that does not affect the optimization objective, we get the loss function of the first part The initial expression: Among them, the subscript NLL indicates that the goal of this part of the loss is to minimize the input normal feature x N The sum of the negative log-likelihoods, N represents the total number of normal samples, D is the set of all normal features, for each negative log-likelihood -log p X (x N ), the first part of the formula represents the mapping result z N The square of the L2 norm of z N The second part of the formula represents the sum of the scaling parameters of the ST conversion subnetwork output in the affine coupling layer contained in the NF, x N ' is a unified reference to the input of the ST conversion subnet; Since the single-scale feature modeling stage uses an independent single-scale flow to model the feature map of each scale, the output feature is expressed as The number of DAPF modules included in each scale is The sum of the scaling parameters of the affine coupling layer output of the j-th single-scale stream is denoted as Among them are: Among them, s(x 1:d ) k and s(y d+1:D ) k They represent the scaling parameters of the two DA-STNet outputs in the k-th DAPF module, respectively. j Indicates the number of DAPF modules used for the feature map of the jth scale; similarly, the multi-scale feature modeling stage uses the multi-scale flow DMFF to model the input of multiple scales simultaneously, and the output is expressed as The output scaling parameters of the two DMFF-STNets in the DMFF block are both multi-scale feature lists, denoted as s_list 1:d and s_list d+1:D , the sum of the scaling parameters S output by its affine coupling layer multi Calculated by the following formula: S multi =∑ j [sum(s_list 1:d ) j +sum(s_list d+1:D ) j ], Among them, sum(s_list) j It means that the j-th feature map in s_list is summed according to the sample dimension, and ∑j means that the sum of each item in s_list is added to obtain the sum of the scaling parameters of the affine coupling layer of this part S multi ; Use the above expression to Rewrite it to get: For optimization objective (2), first calculate the maximum negative log-likelihood of normal samples in the batch, denoted as -log p max , and then set the optimization boundary B of the pseudo-anomaly samples in the batch to: Among them, τ is the minimum distance between the negative log-likelihood of the preset normal features and pseudo-abnormal features; correspondingly, the second part of the optimization objective Calculated as follows: The subscript CLO indicates that the goal of this part of the loss is to optimize the collaborative negative log-likelihood of normal features and pseudo-anomaly features until the distance between the two is greater than τ, N represents the total number of pseudo-anomaly samples, and A is the set of all pseudo-anomaly features; then the total loss function is obtained by the following formula Here, λ is the weight for adding the two losses together.
6. The industrial image anomaly detection method according to claim 5, characterized in that: The abnormality score calculation of the model output described in step S4 is as follows: (1) Pixel-level anomaly score calculation: First, calculate the log-likelihood of the mapping results at each scale, take the average according to the channel dimension, and then upsample it to the input image resolution in terms of spatial resolution. Then, process the upsampled results and use the likelihood in the range of (0,1). Perform multi-scale aggregation and use the accumulation and multiplication method to obtain; Finally, the pixel-level anomaly score map S is generated based on the aggregation results in the following way. loc : S loc =S add =max(P add )-P add , (2) Image-level anomaly score calculation. The specific calculation formula is as follows: S mul =max(P mul )-P mul , First, multiply the result P by multi-scale likelihood mul Calculate the anomaly score map S mul , and then through the abnormal score map S mul The average of the first K pixels with the largest anomaly scores in the image is the final image-level anomaly score s det .
Citation Information
Cited By
Cable surface anomaly detection method and system based on invariant features
CN121661461A
A cable surface anomaly detection method and system based on invariant features
CN121661461B