A multi-class anomaly detection method based on continuous category hints

CN122473197BActive Publication Date: 2026-09-08CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610976495.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-08
Estimated Expiration
2046-07-02

AI Technical Summary

Technical Problem

[0005]基于此,本发明提供了一种基于连续类别提示的多类异常检测方法,以解决背景技术中所提到现有统一多类别异常检测模型中存在的特征分布混叠、多中心建模边界不稳定以及微小缺陷易被背景信息掩盖等问题

Benefits of technology

[0062] As can be seen from the above technical solution, the multi-class anomaly detection method based on continuous category prompts proposed in this invention has the following beneficial effects:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122473197B_ABST
    Figure CN122473197B_ABST
Patent Text Reader

Abstract

The application provides a multi-class anomaly detection method based on continuous category prompts, which comprises the following steps: constructing a multi-class anomaly detection model, which comprises a pre-trained backbone network, a continuous category prompt module, an orthogonal residual decoupling module, a parallel flow module, a context-aware fusion flow module and an anomaly scoring module; extracting multi-scale stage features of an input image through the pre-trained backbone network; generating a continuous category prompt vector based on random binary masks and contrast learning through the continuous category prompt module; orthogonally decomposing shallow stage features in the direction of the continuous category prompt vector through the orthogonal residual decoupling module to obtain common features and residual features, and amplifying the residual; performing reversible transformation under the condition constraint through the parallel flow module and the context-aware fusion flow module to output a multi-scale log-likelihood map; and performing image-level anomaly detection and pixel-level anomaly positioning through the anomaly scoring module. The application realizes efficient anomaly detection and accurate positioning of multi-class industrial products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial anomaly detection technology, and in particular to a multi-class anomaly detection method based on continuous category prompts. Background Technology

[0002] Current unsupervised industrial anomaly detection technologies typically identify anomalous regions that deviate from the feature distribution of normal samples. In actual industrial production lines, it is usually necessary to inspect multiple different categories of products simultaneously. Existing technologies mainly adopt an independent modeling paradigm of "one model per category," that is, training an independent anomaly detection model for each category of products. As the number of industrial inspection categories increases, the number of model parameters and storage overhead increases significantly, resulting in excessively high training and hardware deployment costs.

[0003] To reduce model redundancy, some technologies attempt to adopt a unified multi-class anomaly detection model. However, different product categories exhibit significant differences in texture structure, morphology, and spatial layout. Directly performing unified modeling in a shared feature space easily leads to mutual interference and aliasing of inter-class features. Furthermore, some existing technologies attempt to introduce Gaussian mixture distributions or multi-center latent spaces to enhance class discriminability, but this multi-center discretization modeling disrupts the continuous mapping properties of the normalized flow, causing similar categories to overlap at distribution boundaries. This makes the model insensitive to weak local boundary anomaly signals, limiting the ability to accurately locate defects in multi-class products. Summary of the Invention

[0004] (a) Technical problems to be solved

[0005] Based on this, the present invention provides a multi-class anomaly detection method based on continuous category prompts to solve the problems mentioned in the background art, such as feature distribution aliasing, unstable multi-center modeling boundaries, and the ease with which small defects are masked by background information in existing unified multi-class anomaly detection models.

[0006] (II) Technical Solution

[0007] To achieve the above objectives, the present invention provides a multi-class anomaly detection method based on continuous category prompts, comprising:

[0008] Step S1: Construct a multi-class anomaly detection model;

[0009] The multi-class anomaly detection model includes: a pre-trained backbone network, a continuous category prompting module, an orthogonal residual decoupling module, a conditional guidance module, a parallel flow module, a context-aware fusion flow module, and an anomaly scoring module;

[0010] Let the industrial image input to the multi-class anomaly detection model be... , Multi-scale stage features at four levels were extracted from the pre-trained backbone network. , , , ; Features of Stage 4 Input the continuous category hint module and output the continuous category hint vector. On the one hand, the characteristics of the first stage are respectively Characteristics of Stage 2 Stage 3 characteristics Input the orthogonal residual decoupling module to obtain multi-scale orthogonal residual decoupling features. , , ;Will , , , Inputting each into the parallel stream module yields the following results: , , , On the other hand, , and Input condition guidance module to obtain condition constraints ;Will As conditions and constraints, respectively , , , The input is fed into the context-aware fusion stream module to obtain the log-likelihood graph. , , , ; log-likelihood plot to Input the data into the anomaly scoring module to obtain the anomaly scoring results;

[0011] Step S2: Train the multi-class anomaly detection model;

[0012] Step S3: Input the industrial image to be detected into the multi-type anomaly detection model to obtain the anomaly score result.

[0013] Specifically, in step S1, the pre-trained backbone network is ConvNeXtV2-Base; the core function of the pre-trained backbone network is to extract multi-scale stage features at four levels, namely the first-stage features. Characteristics of Stage 2 Stage 3 characteristics Characteristics of Stage 4 ; It has a high resolution and contains more detailed textures; It has a low resolution but contains more global semantic information.

[0014] Specifically, in step S1, when describing the continuous category prompt module, let the input feature of the module be... Randomly generate a binary mask matrix Then the features after random masking It is expressed as follows:

[0015]

[0016] in, This indicates element-wise multiplication; This represents a randomly generated binary mask matrix with a mask ratio of . In other words, for The elements in the array randomly take values ​​of 0 or 1, and the proportion of values ​​that are 0 is . ;

[0017] Will The input is fed into two different branches, constructing two semantic views of the same image sample; one branch directly addresses... After global pooling, the input suggestion generator produces anchored suggestion vectors. The other branch is... After applying a random drop operation, the result is , After global pooling and a cue generator, the enhanced cue vector is obtained. As shown below:

[0018]

[0019]

[0020] in, This indicates a random discard operation, that is, randomly discarding the feature matrix. Set some elements to 0; Indicates global pooling; The prompt generator is composed of a multi-layered perceptron.

[0021] Then, the anchor cue vector is anchored using the L2 norm. and enhanced cue vectors Projecting onto the unit hypersphere to achieve vector normalization, as shown below:

[0022]

[0023] To enhance semantic separability between different categories, the InfoNCE contrastive learning loss is used to optimize the cue vectors; for the same sample, the normalized anchored cue vectors are... and enhanced cue vectors Positive samples are treated as samples to be brought closer together, while other samples in the same batch are treated as negative samples to be pushed away, thus adaptively forming a cue clustering structure that is compact within classes and separated between classes in the continuous semantic space; the corresponding loss function is:

[0024]

[0025] in, This represents BatchSize, which is the number of samples in a batch. Indicates the temperature coefficient; and These represent the first and second items in the same batch. Anchored cue vector and augmented cue vector for each sample; Indicates the first in the same batch Enhanced cue vectors for each sample;

[0026] Finally, the normalized anchor cue vector As the output of the continuous category hint module, i.e., the continuous category hint vector .

[0027] Specifically, in step S1, when describing the orthogonal residual decoupling module, the input characteristics of the orthogonal residual decoupling module are assumed to be... ,Will In continuous category hint vector Orthogonal decomposition is performed along the direction to obtain the input features. In the prompt vector Projected length in the direction The projection length is used to measure the similarity between the input feature and the overall normal semantics of the current category, as shown below:

[0028]

[0029] in, Indicates the inner product;

[0030] Based on the projection length, Decomposed into parallel components and vertical components Parallel components are common features, and vertical components are residual features, as shown below:

[0031]

[0032]

[0033] This indicates common features consistent with the global normal sample category indication; This indicates the abnormal features that deviate from the normal pattern of the current category;

[0034] To enhance the model's sensitivity to anomalies, a residual amplification module is designed. This module sequentially comprises a first convolution (Conv), normalized batch normalization (BN), a GELU activation function, and a second convolution (Conv). The residual amplification module amplifies the residual features. Feature extraction and enhancement are performed to obtain residual amplification features. Then, common features With residual amplification characteristics The features are concatenated along the channel dimension and then dimensionality reduced using a feature fusion module to obtain the reconstructed decoupled features. The feature fusion module includes, in sequence, convolutional (Conv), normalized batch normalization (BN), and ReLU activation functions.

[0035] Specifically, in step S1, the parallel flow module includes four flow models. - ,Will Input first-order model get ,Will Input second-flow model get ,Will Input third-flow model get ,Will Input fourth-flow model get ;

[0036] The flow model maps a simple distribution to a complex probability distribution through a series of reversible nonlinear transformations. Let its reversible nonlinear transformations be... As shown below:

[0037]

[0038] in, Represents the flow model; This represents the input to the flow model; This represents the output of the flow model;

[0039] First-class model middle, In the second-flow model middle, In the third-flow model middle, In the fourth-flow model middle, .

[0040] Specifically, in step S1, the conditional guidance module has two branches. In one branch, the input continuous category cue vector is... Obtained through a multilayer perceptron (MLP) In another branch, the input shallow detail features are first processed. and After spatial alignment, the pieces are stitched together. Then the obtained splicing features Inputting into a multilayer perceptron (MLP) yields Then, and By splicing and merging the data, we can obtain global and local prompts and constraints. .

[0041] Specifically, in step S1, in the context-aware fusion stream module, firstly, the output of the parallel stream module is... , , After average pooling downsampling, and After achieving the same resolution, then with By splicing the features together, we can obtain the spliced ​​features. As shown below:

[0042]

[0043] in, Indicates average pooling; Indicates splicing;

[0044] Secondly, after a convolution... The channel was reduced to one-quarter. After another convolution, To restore to the original size, that is The dimensions are obtained. ;

[0045] Then, Split to obtain , , , It is worth noting that the splitting here corresponds to the splicing mentioned earlier, that is to say, The splitting point and The joints correspond to each other;

[0046] Finally, respectively , , Upsample to the original corresponding size, that is, respectively... , , Upsampling to and , , Same size, output , , ; and will Directly used as output ;

[0047] For the context-aware fusion stream module, from the input , , , To the output , , , The intermediate process is viewed as a reversible nonlinear transformation similar to a flow model, and here we use... This indicates that the conditional fusion flow model performs a reversible nonlinear transformation under specific constraints, i.e., for... Add condition constraints The conditional fusion flow model is obtained as follows:

[0048]

[0049] in, This represents a conditionally invertible nonlinear transformation from input to output in a context-aware fusion flow module, i.e., a conditional fusion flow model. This represents the input to the context-aware fusion stream module, where ; This indicates the corresponding output of the context-aware fusion stream module.

[0050] Specifically, in step S1, the anomaly scoring module is divided into image-level anomaly detection and pixel-level anomaly localization. The result of image-level anomaly detection is used to determine whether there are anomalies in the industrial image, and the result of pixel-level anomaly localization is used to output the anomaly localization area of ​​the industrial image.

[0051] Regarding image-level anomaly detection, firstly, for the log-likelihood plot output by the conditional fusion flow model... , , , Perform weighted aggregation to obtain the aggregated log-likelihood plot. As shown below:

[0052]

[0053] in, Represents the corresponding weights of the multi-scale log-likelihood plot;

[0054] Then, using the Top-K strategy, the aggregated log-likelihood plot is obtained. The K pixels with the strongest responses are selected, and their average value is calculated as the final image-level anomaly score, as shown below:

[0055]

[0056] in, express The probability density of the top K pixels with the highest median values;

[0057] In pixel-level anomaly localization, due to the low spatial resolution of high-level features, boundary diffusion and halo response are prone to occur during upsampling, affecting the anomaly localization accuracy. Therefore, weight suppression is applied to the highest-level features, resulting in a higher pixel-level localization score. As shown below:

[0058]

[0059] in, This represents the multi-scale probability density map output by the conditional fusion flow model, where ,include: , , , ; This represents the corresponding weights of the multi-scale probability density map.

[0060] Specifically, in step S1, the weights of the log-likelihood plots at each scale are the same, i.e. .

[0061] (III) Beneficial Effects

[0062] As can be seen from the above technical solution, the multi-class anomaly detection method based on continuous category prompts proposed in this invention has the following beneficial effects:

[0063] 1. This invention eliminates the need to train a separate model for each product category. A single model can be used to detect multiple different categories of industrial products simultaneously, significantly reducing the number of model parameters and storage overhead. It also overcomes the problem of linear cost growth when the number of categories increases, which is a problem of the traditional "one model per category" paradigm.

[0064] 2. This invention uses InfoNCE contrastive learning to optimize continuous category hints, which adaptively forms a compact intra-class and separated inter-class hint clustering structure in a unified feature space. This avoids the boundary instability problems caused by discrete category encoding and Gaussian mixture centers, enabling subsequent flow models to accurately model complex multi-class distributions in a continuous semantic space.

[0065] 3. This invention uses the residual decoupling mechanism of spatial orthogonal projection to decompose local features (i.e. stage features) into common components of categories and abnormal residual components. It also enhances abnormal signals through a residual amplification module, effectively removing interference from normal background information and making the model more sensitive to minor defects such as small scratches and local damage.

[0066] 4. This invention achieves excellent performance in both image-level detection and pixel-level localization tasks through a multi-scale conditional fusion flow model and a differentiated anomaly scoring strategy. Attached Figure Description

[0067] The features and advantages of the invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the invention in any way. In the drawings:

[0068] Figure 1 This is a schematic diagram of the structure of the multi-type anomaly detection model of the present invention;

[0069] Figure 2 This is a schematic diagram of the continuous category prompt module of the present invention;

[0070] Figure 3 This is a schematic diagram of the orthogonal residual decoupling module of the present invention;

[0071] Figure 4 This is a schematic diagram of the pixel-level anomaly localization results of the multi-class anomaly detection model of the present invention in five categories (bottle, cable, capsule, carpet, and grid) in the test set.

[0072] Figure 5 This is a schematic diagram of the pixel-level anomaly localization results of the multi-class anomaly detection model of the present invention in five categories (hazelnut, leather, metal nut, pill, and screw) in the test set.

[0073] Figure 6 This is a schematic diagram of the pixel-level anomaly localization results of the multi-class anomaly detection model of the present invention in five categories (tile, toothbrush, transistor, wood, and zipper) in the test set. Detailed Implementation

[0074] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0075] This invention provides a multi-class anomaly detection method based on continuous category prompts, including:

[0076] Step S1: Construct a multi-class anomaly detection model;

[0077] like Figure 1 As shown, the multi-class anomaly detection model includes: a pre-trained backbone network, a continuous category prompting module, an orthogonal residual decoupling module, a conditional guidance module, a parallel flow module, a context-aware fusion flow module, and an anomaly scoring module.

[0078] Let the industrial image input to the multi-class anomaly detection model be... , Multi-scale stage features at four levels are extracted using a pre-trained backbone network (such as ConvNeXtV2-Base). , , , ; Features of Stage 4 Input the continuous category hint module and output the continuous category hint vector. On the one hand, the characteristics of the first stage are respectively Characteristics of Stage 2 Stage 3 characteristics Input the orthogonal residual decoupling module to obtain multi-scale orthogonal residual decoupling features. , , ;Will , , , Inputting each into the parallel stream module yields the following results: , , , On the other hand, , and Input condition guidance module to obtain condition constraints ;Will As conditions and constraints, respectively , , , The input is fed into the context-aware fusion stream module to obtain the log-likelihood graph. , , , ; log-likelihood plot to Input the anomaly scoring module to obtain the anomaly scoring results.

[0079] The core function of the pre-trained backbone network is to extract multi-scale stage features across four levels, namely, stage 1 features. Characteristics of Stage 2 Stage 3 characteristics Characteristics of Stage 4 ; It has a high resolution and contains more detailed textures; With low resolution, these multi-scale feature maps, ranging from fine to coarse, contain more global semantic information, providing necessary, high-quality feature input for subsequent modules. In this embodiment, the input image of the model... The size is , , , , The sizes are respectively , , , .

[0080] When describing the consecutive category hint module, such as Figure 2 As shown, let the input features of this module be... Randomly generate a binary mask matrix Then the features after random masking It is expressed as follows:

[0081]

[0082] in, This indicates element-wise multiplication; This represents a randomly generated binary mask matrix with a mask ratio of . In other words, for The elements in the array randomly take values ​​of 0 or 1, and the proportion of values ​​that are 0 is . Random binary masks can effectively suppress the model's direct memorization of local texture details, allowing the subsequent cue generator to focus more on global category information.

[0083] Will The input is fed into two different branches, constructing two semantic views of the same image sample. One branch directly addresses... After global pooling, the input suggestion generator produces anchored suggestion vectors. The other branch is... After applying a random drop operation, the result is , After global pooling and a cue generator, the enhanced cue vector is obtained. As shown below:

[0084]

[0085]

[0086] in, This indicates a random discard operation, that is, randomly discarding the feature matrix. Set some elements to 0; Indicates global pooling; The prompt generator is composed of a multi-layered perceptron.

[0087] Then, the anchor cue vector is anchored using the L2 norm. and enhanced cue vectors Projecting onto the unit hypersphere to achieve vector normalization, as shown below:

[0088]

[0089] To enhance semantic separability between different categories, the InfoNCE contrastive learning loss is used to optimize the cue vectors. For the same sample, the normalized anchored cue vector is... and enhanced cue vectors Positive samples are treated as samples to be brought closer together, while other samples in the same batch are treated as negative samples to be pushed away, thus adaptively forming a cue clustering structure that is compact within classes and separated between classes in the continuous semantic space; the corresponding loss function is:

[0090]

[0091] in, This represents BatchSize, which is the number of samples in a batch. Indicates the temperature coefficient; and These represent the first and second items in the same batch. Anchored cue vector and augmented cue vector for each sample; Indicates the first in the same batch Augmentation cue vectors for each sample.

[0092] Finally, the normalized anchor cue vector As the output of the continuous category hint module, i.e., the continuous category hint vector .

[0093] The continuous category hint module dynamically generates global continuous category hint vectors from input features through random binary masks, multi-view semantic construction and hint vector generation, normalization, and contrastive learning constraints, constructing stable category hints in a unified feature space. Unlike traditional discrete category encoding and Gaussian mixture centers, continuous category hints do not require predefined fixed category centers or discrete codebooks. Instead, they adaptively learn continuous semantic relationships between categories through a feature-driven approach, forming semantically robust class-conditional representations and enhancing the model's unified modeling capability under complex category differences and texture variations.

[0094] When describing the orthogonal residual decoupling module, such as Figure 3 As shown, let the input characteristics of the orthogonal residual decoupling module be... ,Will In continuous category hint vector Orthogonal decomposition is performed along the direction to obtain the input features. In the prompt vector Projected length in the direction The projection length is used to measure the similarity between the input feature and the overall normal semantics of the current category, as shown below:

[0095]

[0096] in, This indicates the inner product.

[0097] Based on the projection length, Decomposed into parallel components (i.e., common features) and vertical components (i.e., residual characteristics), as shown below:

[0098]

[0099]

[0100] This indicates common features consistent with the global normal sample category indication; This indicates the abnormal feature portion that deviates from the normal pattern of the current category.

[0101] To enhance the model's sensitivity to anomalies, a residual amplification module is designed. This module sequentially comprises a first convolution (Conv), normalization (BN), a GELU activation function, and a second convolution (Conv). The residual amplification module amplifies the residual features. Feature extraction and enhancement are performed to obtain residual amplification features. Then, common features With residual amplification characteristics The features are concatenated along the channel dimension and then dimensionality reduced using a feature fusion module to obtain the reconstructed decoupled features. The feature fusion module includes convolution (Conv), normalization (BN), and ReLU activation functions.

[0102] By orthogonally projecting the feature space, the input stage features are decomposed into common feature components with consistent categories and abnormal residual feature components that deviate from the normal pattern. Abnormal residuals are amplified in a targeted manner to achieve decoupled expression of normal background information and potential abnormal information.

[0103] The parallel stream module includes 4 stream models ( - ),Will Input first-order model get ,Will Input second-flow model get ,Will Input third-flow model get ,Will Input fourth-flow model get .

[0104] The flow model, a current technology, maps a simple distribution to a complex probability distribution through a series of reversible nonlinear transformations. Let its reversible nonlinear transformations be... As shown below:

[0105]

[0106] in, Represents the flow model; This represents the input to the flow model; This represents the output of the flow model;

[0107] First-class model middle, In the second-flow model middle, In the third-flow model middle, In the fourth-flow model middle, .

[0108] For the features in stages 1-3, the orthogonal residual decoupling features guided by the prompts are used as input to eliminate background interference; for the features in stage 4, the continuous category prompt vector P is used as input to guide the flow model to learn the continuous distribution relationship of different categories in the semantic space.

[0109] The conditional guidance module has two branches. In one branch, the input is a continuous category cue vector. Obtained through a multilayer perceptron (MLP) In another branch, the input shallow detail features are first processed. and After spatial alignment, the pieces are stitched together. Then the obtained splicing features Inputting into a multilayer perceptron (MLP) yields Then, and By splicing and merging the data, we can obtain global and local prompts and constraints. .

[0110] In the context-aware fusion streaming module, firstly, the output of the parallel streaming module is... , , After average pooling downsampling, and After achieving the same resolution, then with By splicing the features together, we can obtain the spliced ​​features. As shown below:

[0111]

[0112] in, Indicates average pooling; This indicates splicing; in this embodiment, the spliced ​​features Size is .

[0113] Secondly, after a convolution... The channel was reduced to one-quarter. ,Right now The size is After another convolution, Restored to its original size (i.e.) Size ),get .

[0114] Then, Split to obtain , , , It is worth noting that the splitting and formula here... The splicing corresponds to the, that is to say The splitting point and The joints correspond to each other.

[0115] Finally, respectively , , Upsample to the original corresponding size, that is, respectively... , , Upsampling to and , , Same size, output , , ; and will Directly used as output In this embodiment, , , , The dimensions are respectively , , , .

[0116] For the context-aware fusion stream module, from the input , , , To the output , , , The intermediate process can be viewed as a reversible nonlinear transformation similar to a flow model, which is used here. This is indicated. The conditional fusion flow model of this invention performs a reversible nonlinear transformation under specific conditional constraints, that is, for... Add condition constraints The conditional fusion flow model is obtained as follows:

[0117]

[0118] in, This represents a conditional invertible nonlinear transformation from input to output in the context-aware fusion flow module (i.e., a conditional fusion flow model). This represents the input to the context-aware fusion stream module; This indicates the corresponding output of the context-aware fusion stream module.

[0119] Due to the global continuous category cue vector Stage characteristics , Both originate from continuous semantic space. The fused stream can achieve continuous fitting of complex distributions of multiple categories under a unified Gaussian distribution center, and has strong category discrimination ability and pixel-level anomaly sensitivity.

[0120] Traditional unsupervised anomaly detection methods typically employ a uniform scale aggregation strategy to perform image-level detection and pixel-level localization tasks. However, features at different scales play different roles in these two types of tasks. High-level semantic features have stronger global structural representation capabilities and are more suitable for describing overall structural anomalies, while shallow texture features retain more edge and local detail information, making them more conducive to the fine-grained localization of minor defects. Therefore, this invention addresses the differences in the roles of features at different scales in the two types of tasks by setting different weights.

[0121] The anomaly scoring module is divided into image-level anomaly detection and pixel-level anomaly localization. The results of image-level anomaly detection are used to determine whether there are anomalies in the industrial image, while the results of pixel-level anomaly localization are used to output the anomaly localization area of ​​the industrial image.

[0122] Regarding image-level anomaly detection, firstly, for the log-likelihood plot output by the conditional fusion flow model... , , , Perform weighted aggregation to obtain the aggregated log-likelihood plot. As shown below:

[0123]

[0124] in, This represents the corresponding weights of the multi-scale log-likelihood plots. In this embodiment, since image-level anomaly detection requires preserving the complete global structural response, the weights of the log-likelihood plots at each scale are the same.

[0125] Then, using the Top-K strategy, the aggregated log-likelihood plot is obtained. The K pixels with the strongest response (the preferred value of K is 3% of all pixels) are averaged as the final image-level anomaly score, as shown below:

[0126]

[0127] in, express The probability density of the top K pixels with the highest median values.

[0128] Compared to directly using the maximum value strategy, Top-K averaging can effectively reduce the instability caused by noise response and improve the overall detection robustness.

[0129] In pixel-level anomaly localization, due to the low spatial resolution of high-level features, boundary diffusion and halo response are easily generated during upsampling, affecting the anomaly localization accuracy. Therefore, this invention performs weight suppression on the highest-level features (stage 4), resulting in a higher pixel-level localization score. As shown below:

[0130]

[0131] in, The multi-scale probability density map representing the output of the conditional fusion flow model includes: , , , ; This represents the corresponding weights of the multi-scale probability density map.

[0132] Step S2: Train the multi-class anomaly detection model;

[0133] The two industry benchmark datasets, MVTec AD and VisA, were divided into training and testing sets according to a certain ratio, and the multi-class anomaly detection model was trained using the training set.

[0134] Step S3: Input the industrial image to be detected into the multi-type anomaly detection model to obtain the anomaly score result.

[0135] In this embodiment, the test set is input into the multi-class anomaly detection model. The test datasets of MVTec AD and VisA achieve accuracy rates of 99.7% and 97.2% respectively in image-level anomaly detection, and 98.8% and 99.0% respectively in pixel-level anomaly localization. Figure 4-6As shown, a visual heatmap of the anomaly localization results for 15 categories in the test set is presented. The results demonstrate that this invention can accurately locate not only minute scratches, stains, and structural damage on the surfaces of products with complex textures (such as carpets, grids, tiles, wood, and leather), but also highly adaptable to component defects or deformations in products with complex three-dimensional shapes and structures (such as cables, transistors, metal nuts, and toothbrushes). The heatmap and the actual mask maintain a very high degree of spatial consistency. The highlighted areas (red-to-yellow transition zone) in the heatmap closely match the actual anomaly boundaries, while the normal areas exhibit a uniform and clean low-response state (dark blue). Even when faced with small scratches on wood and leather, or through cracks in tiles, this invention still generates a focused and strong anomaly response, without response diffusion or large-area false alarms caused by complex background features. This fully demonstrates that the orthogonal residual decoupling and continuous category hinting mechanism in this invention can effectively isolate normal common background and has a good ability to locate anomalies under the condition of multi-class unified modeling.

[0136] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-class anomaly detection method based on continuous category prompts, characterized in that, include: Step S1: Construct a multi-class anomaly detection model; The multi-class anomaly detection model includes: a pre-trained backbone network, a continuous category prompting module, an orthogonal residual decoupling module, a conditional guidance module, a parallel flow module, a context-aware fusion flow module, and an anomaly scoring module; Let the industrial image input to the multi-class anomaly detection model be... , Multi-scale stage features at four levels were extracted from the pre-trained backbone network. , , , ; Features of Stage 4 Input the continuous category hint module and output the continuous category hint vector. On the one hand, the characteristics of the first stage are respectively Characteristics of Stage 2 Stage 3 characteristics Input the orthogonal residual decoupling module to obtain multi-scale orthogonal residual decoupling features. , , ;Will , , , Inputting each into the parallel stream module yields the following results: , , , On the other hand, , and Input condition guidance module to obtain condition constraints ;Will As conditions and constraints, respectively , , , The input is fed into the context-aware fusion stream module to obtain the log-likelihood graph. , , , Log-likelihood plot to Input the data into the anomaly scoring module to obtain the anomaly scoring results; When describing the continuous category hint module, let the input features of the module be: Randomly generate a binary mask matrix Then the features after random masking It is expressed as follows: in, This indicates element-wise multiplication; This represents a randomly generated binary mask matrix with a mask ratio of . ;right The elements in the array randomly take values ​​of 0 or 1, and the proportion of values ​​that are 0 is . ; Will The input is fed into two different branches, constructing two semantic views of the same image sample; one branch directly addresses... After global pooling, the input suggestion generator produces anchored suggestion vectors. The other branch is... After applying a random drop operation, the result is , After global pooling and a cue generator, the enhanced cue vector is obtained. As shown below: in, This indicates a random discard operation, that is, randomly discarding the feature matrix. Set some elements to 0; Indicates global pooling; The prompt generator is composed of a multi-layered perceptron. Then, the anchor cue vector is anchored using the L2 norm. and enhanced cue vectors Projecting onto the unit hypersphere to achieve vector normalization, as shown below: To enhance semantic separability between different categories, the InfoNCE contrastive learning loss is used to optimize the cue vectors; for the same sample, the normalized anchored cue vectors are... and enhanced cue vectors Positive samples are treated as samples to be brought closer together, while other samples in the same batch are treated as negative samples to be pushed away, thus adaptively forming a cue clustering structure that is compact within classes and separated between classes in the continuous semantic space; the corresponding loss function is: in, This represents BatchSize, which is the number of samples in a batch. Indicates the temperature coefficient; and These represent the first and second items in the same batch. Anchored cue vector and augmented cue vector for each sample; Indicates the first in the same batch Enhanced cue vectors for each sample; Finally, the normalized anchor cue vector As the output of the continuous category hint module, i.e., the continuous category hint vector ; When describing the orthogonal residual decoupling module, let the input characteristics of the orthogonal residual decoupling module be... ,Will In continuous category hint vector Orthogonal decomposition is performed along the direction to obtain the input features. In the prompt vector Projected length in the direction The projection length is used to measure the similarity between the input feature and the overall normal semantics of the current category, as shown below: in, Indicates the inner product; Based on the projection length, Decomposed into parallel components and vertical components Parallel components are common features, and vertical components are residual features, as shown below: This indicates common features consistent with the global normal sample category indication; This indicates the abnormal features that deviate from the normal pattern of the current category; To enhance the model's sensitivity to anomalies, a residual amplification module is designed. This module sequentially comprises a first convolution (Conv), normalized batch normalization (BN), a GELU activation function, and a second convolution (Conv). The residual amplification module amplifies the residual features. Feature extraction and enhancement are performed to obtain residual amplification features. Then, common features With residual amplification characteristics The features are concatenated along the channel dimension and then dimensionality reduced using a feature fusion module to obtain the reconstructed decoupled features. The feature fusion module includes, in sequence, convolutional (Conv), normalized batch normalization (BN), and ReLU activation functions. Step S2: Train the multi-class anomaly detection model; Step S3: Input the industrial image to be detected into the multi-type anomaly detection model to obtain the anomaly score result.

2. The method according to claim 1, characterized in that, In step S1, the pre-trained backbone network is ConvNeXtV2-Base; the core function of the pre-trained backbone network is to extract multi-scale stage features at four levels, namely the first-stage features. Characteristics of Stage 2 Stage 3 characteristics Characteristics of Stage 4 ; It has a high resolution and contains more detailed textures; It has a low resolution but contains more global semantic information.

3. The method according to claim 2, characterized in that, In step S1, the parallel flow module includes four flow models. - ,Will Input first-order model get ,Will Input second-flow model get ,Will Input third-flow model get ,Will Input fourth-flow model get ; The flow model maps a simple distribution to a complex probability distribution through a series of reversible nonlinear transformations. Let its reversible nonlinear transformations be... As shown below: in, Represents a flow model; This represents the input to the flow model; This represents the output of the flow model; First-class model middle, In the second-flow model middle, In the third-flow model middle, In the fourth-flow model middle, .

4. The method according to claim 3, characterized in that, In step S1, the conditional guidance module has two branches. In one branch, the input continuous category cue vector is... After being processed by a multilayer perceptron (MLP) In another branch, the input shallow detail features are first processed. and After spatial alignment, the pieces are stitched together. Then the obtained splicing features Inputting into a multilayer perceptron (MLP) yields Then, and By splicing and merging the data, we can obtain global and local prompts and constraints. .

5. The method according to claim 4, characterized in that, In step S1, within the context-aware fusion streaming module, firstly, the output of the parallel streaming module is... , , After average pooling downsampling, and After achieving the same resolution, then with By splicing the features together, we can obtain the spliced ​​features. As shown below: in, Indicates average pooling; Indicates splicing; Secondly, after a convolution... The channel was reduced to one-quarter. After another convolution, To restore to the original size, that is The dimensions are obtained. ; Then, Split to obtain , , , The splitting here corresponds to the splicing mentioned earlier. The splitting point and The joints correspond to each other; Finally, respectively , , Upsample to the original corresponding size, that is, respectively... , , Upsampling to and , , Same size, output , , And will Directly used as output ; For the context-aware fusion stream module, from the input , , , To the output , , , The intermediate process is viewed as a reversible nonlinear transformation similar to a flow model, and here we use... This indicates that the conditional fusion flow model performs a reversible nonlinear transformation under specific constraints, i.e., for... Add condition constraints The conditional fusion flow model is obtained as follows: in, This represents a conditionally invertible nonlinear transformation from input to output in a context-aware fusion flow module, i.e., a conditional fusion flow model. This represents the input to the context-aware fusion stream module, where ; This indicates the corresponding output of the context-aware fusion stream module.

6. The method according to claim 5, characterized in that, In step S1, the anomaly scoring module is divided into image-level anomaly detection and pixel-level anomaly localization. The result of image-level anomaly detection is used to determine whether there are anomalies in the industrial image, and the result of pixel-level anomaly localization is used to output the anomaly localization area of ​​the industrial image. Regarding image-level anomaly detection, firstly, for the log-likelihood plot output by the conditional fusion flow model... , , , Perform weighted aggregation to obtain the aggregated log-likelihood plot. As shown below: in, Represents the corresponding weights of the multi-scale log-likelihood plot; Then, using the Top-K strategy, the aggregated log-likelihood plot is obtained. The K pixels with the strongest responses are selected, and their average value is calculated as the final image-level anomaly score, as shown below: in, express The probability density of the top K pixels with the highest median values; In pixel-level anomaly localization, due to the low spatial resolution of high-level features, boundary diffusion and halo response are prone to occur during upsampling, affecting the anomaly localization accuracy. Therefore, weight suppression is applied to the highest-level features, resulting in a higher pixel-level localization score. As shown below: in, This represents the multi-scale probability density map output by the conditional fusion flow model, where ,include: , , , ; This represents the corresponding weights of the multi-scale probability density map.

7. The method according to claim 6, characterized in that, In step S1, the weights of the log-likelihood plots at each scale are the same, that is... .

Citation Information

Patent Citations

  • Multi-category industrial image anomaly detection method and device based on three-tower knowledge distillation architecture

    CN119625412A

  • Adaptive anomaly detection method based on feature fusion and screening

    CN122176475A