Method for constructing zero-shot defect detection model in industrial field

CN122841318APending Publication Date: 2026-09-29XINJIANG AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611008893.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-11-04
Filing Date
2026-07-08
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0007]为解决现有方法在跨模态信息融合、语义理解以及全局异常特征提取不足、频域特征的敏感性有限可能导致边缘区域被误分类为异常的问题,本发明提出一种工业领域零样本缺陷检测模型(以下简称SAFER-CLIP)

Benefits of technology

(1)设计了基于流形假设的自监督重建模块。该模块的核心机制在于建立重建损失机制的数学表达式,定义流形上正常数据的高保真重建,从而为异常检测提供清晰的判别依据。在此基础上,进一步引入了动态自适应的异常阈值机制,该机制能够根据样本内容自主调整,有效提升了异常检测的精准度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841318A_ABST
    Figure CN122841318A_ABST
Patent Text Reader

Abstract

This invention relates to the fields of computer vision and anomaly detection technology, and in particular to a method for constructing a zero-shot defect detection model for industrial applications. The defect / anomaly detection model includes: a self-supervised reconstruction module, a frequency-aware enhancement module, a cascaded MT architecture, a text encoder, and a total loss module. This invention addresses the problems of existing methods, such as insufficient cross-modal information fusion, semantic understanding, and global anomaly feature extraction, and the limited sensitivity of frequency domain features, which may lead to edge regions being misclassified as anomalies.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to invention patent application No. 202511602564.7, filed on November 4, 2025. Technical Field

[0002] This invention relates to the fields of computer vision and anomaly detection technology, and in particular to a method for constructing a zero-shot defect detection model for industrial applications. Background Technology

[0003] Thanks to significant advancements in computer vision, anomaly detection technology can identify anomalous samples or data deviating from expected patterns through computational methods. This key technology has played a crucial role in various fields, including industrial inspection and medical diagnosis. Existing anomaly detection methods are primarily based on unsupervised and semi-supervised learning, but their practical application faces severe challenges—the models are highly dependent on high-quality, high-standard training data. In many real-world scenarios, the inherent scarcity of anomalous samples and the resulting high acquisition costs make it difficult to meet stringent data standards. This inherent difficulty in data acquisition not only severely restricts the model's learning and recognition capabilities but also hinders its widespread deployment and accuracy maintenance.

[0004] In recent years, zero-shot anomaly detection methods have achieved direct image anomaly detection without the need for training on specific anomaly samples by leveraging the rich semantic understanding capabilities of pre-trained models. However, WinCLIP, a key visual language model for zero-shot anomaly detection, struggles to recognize diverse anomaly details and cannot effectively detect anomalies of different scales and types. AnomalyCLIP employs object-independent text prompts for cross-domain adaptation, enhancing anomaly focusing and generalization capabilities by capturing both normal and anomalous features. However, its independent image-text mapping mechanism limits its deep understanding of image features, and its sensitivity to specific prompt words may lead to overfitting.

[0005] Anomaly detection methods are traditionally divided into unsupervised anomaly detection (UAD) and semi-supervised anomaly detection. With the advent of deep learning, various neural network-based UAD methods have emerged. Among them, autoencoder (AE)-based anomaly detection methods aim to learn accurate reconstructions of normal samples, but their performance is highly sensitive to the purity of the training data. Generative Adversarial Network (GAN)-based anomaly detection methods utilize a large amount of normal data to train the GAN model, enabling its generator to simulate similar outputs. This method typically requires a large number of normal data samples. Furthermore, methods such as density estimation and self-supervised contrastive learning also rely on learning from unlabeled normal data. In contrast, deep semi-supervised anomaly detection methods utilize both normal data and labeled anomaly data during training. By introducing labeled anomaly samples, these methods can learn more compact and accurate decision boundaries for normal data, often achieving anomaly detection performance superior to UAD methods. However, supervised anomaly detection (SSAD) has even higher data requirements, needing to acquire and label a sufficient number of anomaly samples, which may be limited by data availability in real-world scenarios. These methods typically exhibit limited ability to detect novel or unseen anomalies, especially when there is a lack of available anomaly instances.

[0006] Zero-shot anomaly detection (ZSAD) identifies unknown anomalies using only normal samples or common knowledge, without requiring anomalous training samples. However, achieving robust generalization typically requires extensive model training. Recent advances in visual-language models (VLMs) have led to the widespread application of pre-trained VLM-based anomaly detection. Notably, the well-known CLIP model is a focal point in anomaly detection, with methods such as CLIP-AD. Despite these efforts, methods like Win-CLIP improve detection accuracy by enhancing CLIP with text annotations, but their limitation lies in relying on image-text pre-training rather than direct anomaly detection training. Similarly, AdaCLIP and VCP-CLIP leverage textual information through various learnable cue words to improve ZSAD performance; however, this dependence on text may hinder the model's generalization ability. CLIP-Mamba was the first to apply Mamba to anomaly detection, where a pre-trained encoder works in conjunction with a novel Mamba-based decoder to effectively capture multi-scale anomaly features. Building upon the success of previous CLIP models, the AnomalyCLIP model further optimizes the fusion of global and local contextual information. However, under the interference of complex backgrounds and high-frequency normal edges or textures in images, the model still faces the problem of difficulty in accurately distinguishing abnormal and normal regions, and its ability to identify anomalies at multiple scales is insufficient. In addition, the independent image-text mapping limits the deep understanding of image features, and sensitivity to specific cues may lead to overfitting. Summary of the Invention

[0007] To address the shortcomings of existing methods in cross-modal information fusion, semantic understanding, and global anomaly feature extraction, as well as the limited sensitivity of frequency domain features which may lead to misclassification of edge regions as anomalies, this invention proposes a zero-shot defect detection model for industrial applications (hereinafter referred to as SAFER-CLIP). Extensive experiments on benchmark datasets demonstrate that the SAFER-CLIP model can effectively fuse cross-modal information and capture frequency domain anomalies with high accuracy.

[0008] This invention proposes a method for constructing a zero-shot defect detection model in the industrial field. First, a visual language model is constructed, comprising: a self-supervised reconstruction module, a frequency-aware enhancement module, a cascaded MT architecture, a text encoder, and a total loss module. The self-supervised reconstruction module reconstructs the input image. The reconstructed image is processed by the cascaded MT architecture, which outputs multi-layer features from shallow to deep layers. The shallowest and deepest features are processed by the frequency-aware enhancement module to obtain enhanced features. The enhanced features output by the frequency-aware enhancement module are fused with the intermediate layer features output by the cascaded MT architecture to obtain reconstructed features. The deepest feature output by the cascaded MT architecture is also output as an image feature. The text encoder processes the set text prompts to generate text features that tend towards the image category. The text prompts include normal text prompts and abnormal text prompts, and the image categories include normal image samples and defective image samples. Normal images are defect-free image samples. The total loss module aligns the reconstructed features, image features, and text features pairwise, calculates the alignment loss, and fuses the three alignment losses to obtain the total loss function. During training, normal images are extracted as training samples and input into the visual language model. After processing, the total loss function is obtained. Backpropagation is then performed based on the total loss function to update the visual language model. Repeat the above steps until the visual language model converges.

[0009] Preferably, the defect detection model first trains a self-supervised reconstruction module on a normal image dataset; then it trains the defect detection model on a learning dataset that contains both normal image samples and defective image samples, and updates the self-supervised reconstruction module, the cascaded MT architecture, the frequency-aware enhancement module, and the text encoder. During the training of the self-supervised reconstruction module, the loss function adopts the alignment loss between the intermediate layer features output by the cascaded MT architecture and the reconstructed features; The loss function used in the overall model training process includes image features. Reconstruction features and text features The total loss is obtained by weighted summation of the alignment losses between each pair of elements.

[0010] Preferably, the loss function used during the overall model training process for:

[0011] in, Indicates the total loss. To rebuild losses Weighting coefficients, reconstruction loss Alignment loss between reconstructed features and intermediate layer features output by the serial MT architecture.

[0012] Preferably, the self-supervised reconstruction module consists of five identical CNN-Mamba feature enhancement modules connected sequentially, and the CNN-Mamba feature enhancement module consists of CNN modules and Mamba modules connected sequentially. The CNN module consists of sequentially connected transposed convolutional layers, the GELU activation function, and layer normalization; the Mamba module includes a sequentially connected first linear projection layer, convolutional layers, a first nonlinear transformation layer, a state-space model SSM, and a third linear projection layer; it also includes a second linear projection layer and a second nonlinear transformation layer; The output of the normalized layer is connected to the first linear projection layer and the second linear projection layer as the output of the CNN module. The output of the second linear projection layer is connected to the input of the second nonlinear transformation layer. The output of the second nonlinear transformation layer is multiplied element-wise with the output of the state-space model SSM and then used as the input of the third linear projection layer. The output of the third linear projection layer is used as the output of the CNN-Mamba feature enhancement module.

[0013] Preferably, the serial MT architecture includes four feature interaction modules with identical structures connected in series; the feature interaction module consists of a Res-Attention module and a Res-Mamba module connected sequentially. The Res-Attention module comprises a sequentially connected 3D Fast Fourier Transform (FFT) layer, a linear projection layer, an attention computation layer, an output projection layer, a first normalization layer, a feedforward network, a second normalization layer, and a dimension reduction layer. The attention computation layer employs multiple parallel single-head attention branches; the outputs of these branches are concatenated and projected to obtain the attention fusion feature. The input features of the Res-Attention module undergo 2D FFT and linear projection, followed by attention head splitting. Each attention head processes its input, and the outputs of these attention heads are concatenated and passed through the output projection layer to obtain the output projection. The output projection is processed by the first normalization layer and then residuals are calculated. These residual features are fed into the feedforward network (FFN) for further processing. The output of the FFN is processed by the second normalization layer and then residuals are calculated again. Dimension reduction of these residual features yields the output of the Res-Attention module. The Res-Mamba module consists of a sequentially connected dimension-flattening layer, a state-space model (SSM), and a third normalization layer. After dimension flattening, the input of the Res-Mamba module is processed by the state-space model (SSM). The output features of the state-space model (SSM) are projected onto the input feature space of the third normalization layer. The normalized output of the third normalization layer is left as a residual with its input features, and then the dimensions are restored to obtain the output of the Res-Mamba module.

[0014] Preferably, the frequency awareness enhancement module divides the input spliced ​​feature frequency bands into low-frequency features, power frequency features, and high-frequency features according to a frequency threshold and unifies the dimensions; then, it performs Fourier transforms on each to obtain low-frequency frequency domain features, power frequency frequency domain features, and high-frequency frequency domain features respectively; then, according to the set frequency domain allocation mechanism, it applies a frequency band mask to each frequency domain feature to obtain the target frequency features of each frequency band; then, the frequency feature fusion enhancement network uses adaptive weights to fuse the target frequency features of each frequency band to generate enhanced features.

[0015] Preferably, the defect detection model consists of a trained visual language model and a detection output module; the detection output module is used to calculate a comprehensive anomaly score by weighted summation of the distances between each pair of reconstructed features, image features, and text features. The image samples to be detected are judged based on the comprehensive anomaly score. Is it a normal image or a defective image? Comprehensive Abnormal Score For reconstruction features Image features and text features Weighted summation of the distances between each pair of pairs.

[0016] Preferably, the detection output module will integrate the anomaly score. With threshold Compare; if the overall anomaly score is less than the threshold Then determine the input image sample to be detected. Normal; otherwise, it is judged as abnormal, that is, there is a defect. threshold Take the 95th percentile of the comprehensive abnormal scores of normal samples.

[0017] Preferably, the threshold , and These are the mean and standard deviation of the combined outlier scores on the normal validation set, respectively. This is an empirical coefficient.

[0018] The present invention proposes a system for constructing a zero-sample defect detection model in the industrial field, comprising a memory and a processor. The memory stores a computer program, and the processor is connected to the memory. The processor is used to execute the computer program to realize the method for constructing the zero-sample defect detection model in the industrial field.

[0019] The advantages of this invention are: (1) A self-supervised reconstruction module based on the manifold hypothesis was designed. The core mechanism of this module is to establish a mathematical expression for the reconstruction loss mechanism and define high-fidelity reconstruction of normal data on the manifold, thereby providing a clear basis for anomaly detection. On this basis, a dynamic adaptive anomaly threshold mechanism is further introduced, which can adjust autonomously according to the sample content, effectively improving the accuracy of anomaly detection.

[0020] (2) The core innovation of the frequency awareness enhancement module lies in the ingenious use of the characteristic of abnormal signals having a surge in energy within a specific frequency band. The module is designed with a multi-band decomposition mechanism to accurately capture anomalies in the frequency domain and to design a mathematical expression for a global modeling mechanism to help the model understand and fuse different frequency characteristics.

[0021] (3) Based on Passevar's identity, the drastic local structural changes caused by anomalies manifest as a significant increase in the energy of high-frequency components. We establish a novel mathematical expression for a frequency-domain mechanism to provide an adaptive multi-level feature fusion strategy. This strategy adaptively weights and fuses features from different encoder stages in the frequency domain, focusing on specific frequency-domain energy changes caused by anomalies, thereby comprehensively and accurately capturing various forms of anomaly patterns. This step is particularly crucial because it enhances the model's ability to distinguish between anomalies and edges in the high-frequency region, effectively preventing misjudgments.

[0022] (4) The serial MT architecture cleverly combines the advantages of two complementary mechanisms, Mamba and Transformer: the Mamba model, with its excellent sequence scanning capability, is good at efficiently capturing long-range dependencies; while the Transformer, with its powerful self-attention mechanism, can effectively model global feature interactions. This unique fusion mechanism, through synergistic effects, not only significantly optimizes the boundary between normal data manifolds and anomalous patterns, thereby greatly enhancing their separability, but also greatly improves the robustness of the model to distribution shifts.

[0023] Compared with the baseline model SAFER-CLIP . Compared to the previous model, after fully integrating the three modules, the model achieved a significant and stable improvement in anomaly detection metrics. This result confirms that the fusion mechanism effectively achieves long-range dependency modeling and selective high-level semantic alignment by capturing more subtle anomaly patterns and enhancing feature representation, thereby significantly improving the model's overall anomaly detection performance and generalization ability.

[0024] (5) The zero-sample defect detection model for the industrial field constructed in this invention has broad applicability and flexibility, significantly improves pixel-level performance, and highlights the clear effectiveness, universality, and scalability of the results of the frequency sensing enhancement module. These important findings, when accumulated, strongly validate our proposed anomaly pattern capture strategy based on frequency domain anomalous energy changes, affirming its outstanding effectiveness, critical importance, and broad applicability.

[0025] (6) To further evaluate the generalization ability of the zero-shot defect detection model for industrial applications constructed in this invention, comprehensive experiments were conducted on six professional medical datasets. The experimental results on these six datasets show that the model demonstrates further improvement over AnomalyCLIP in various average metrics, consistently achieving good performance. This model utilizes a hybrid architecture of serial MT architecture, which not only effectively captures long-range dependencies but also promotes selective high-level semantic alignment. Simultaneously, by integrating a frequency-aware enhancement module, which utilizes multi-frequency decomposition and adaptive feature fusion techniques, the model's ability to distinguish between anomalies and edges is significantly enhanced. These advancements enable the model to adapt more robustly and accurately to medical image data affected by various complex features, thus demonstrating its powerful generalization ability in identifying complex and diverse features.

[0026] (7) The zero-shot defect detection model for the industrial field constructed in this invention enhances anomaly sensitivity through residual reconstruction achieved by a self-supervised reconstruction mechanism based on the manifold assumption, and improves the ability to identify boundary and edge anomalies by utilizing a frequency-aware enhancement module. This enables it to achieve the most competitive performance in the area under the receiver operating characteristic curve (AUROC) metric across datasets, demonstrating powerful image-level anomaly detection capabilities. This advanced fusion strategy significantly improves the accuracy and stability of the model constructed in this invention for anomaly detection in more challenging industrial environments. For high-texture mesh images, the model constructed in this invention can accurately identify real defects while maintaining the cleanliness and normality of the background area, indicating that the designed self-supervised reconstruction mechanism and frequency-aware enhancement module fusion strategy can effectively distinguish between anomalies and significant high-frequency features. Attached Figure Description

[0027] Figure 1(a) is a module connection diagram of SAFER-CLIP; Figure 1(b) shows the network topology of SAFER-CLIP; Figure 2 The module connection diagram for the self-supervised reconstruction module; Figure 3 This is a module connection diagram for the CNN-Mamba feature enhancement module; Figure 4 This is a module connection diagram for a serial MT architecture; Figure 5 The module connection graph for the Res-Attention and Res-Mamba feature interaction modules; Figure 6 This is a module connection diagram for the frequency sensing enhancement module; Figure 7 A visualization of continuous anomaly detection; Figure 8(a) shows the visualization results of anomaly detection in a cable; Figure 8(b) shows the visualization results of abnormality detection in the capsule (shell); Figure 8(c) shows the visualization results of anomaly detection in the grid; Figure 8(d) shows the visualization results of the anomaly detection on the bottle; Figure 9 This is a flowchart illustrating the method for constructing a defect detection model. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] This embodiment proposes a method for constructing a zero-shot defect detection model in the industrial field. First, a defect detection model SAFER-CLIP is built, and then zero-shot training and learning are performed.

[0030] Figure 1(a) shows the defect detection model SAFER-CLIP, which includes: a self-supervised reconstruction module, a serial MT architecture, a frequency-aware enhancement module, a text encoder, and a detection output module.

[0031] The self-supervised reconstruction module is used to reconstruct the input image. The reconstructed image is processed by a cascaded MT architecture, which outputs multiple layers of features from shallow to deep. The shallowest and deepest features are then processed by a frequency-aware enhancement module to obtain enhanced features. The enhanced features output by the frequency-aware enhancement module are fused with the intermediate layer features output by the cascaded MT architecture to obtain the reconstructed features. The fusion method can employ dimensional overlay; the deepest features output by the serial MT architecture are also used as image features. Output; The text encoder processes the set text prompts to generate text features that tend towards the image category. The detection output module is used to analyze the reconstructed features. Image features and text features The overall anomaly score is calculated by summing the weighted distances between each pair of points. Based on the comprehensive anomaly score and the set threshold The comparison results determine the image sample to be detected. Is it normal? Specifically, if the overall anomaly score is less than the threshold... Then determine the input image sample to be detected. Normal; otherwise, it is judged as abnormal, that is, there is a defect.

[0032] Specifically, the formula for calculating the comprehensive anomaly score is as follows:

[0033] in, This indicates the text features corresponding to normal text prompts. Cosine distance can be used:

[0034] The method for judging image samples is as follows:

[0035] threshold This can be determined statistically from the normal validation set, for example, by taking the 95th percentile of the combined anomaly scores of the normal samples:

[0036] in, This represents the normal validation set, which contains only normal image samples.

[0037] threshold Alternatively, the mean and standard deviation can be used to determine this:

[0038] in, and These are the mean and standard deviation of the combined outlier scores on the normal validation set, respectively. This is an empirical coefficient.

[0039] The self-supervised reconstruction module consists of five identical CNN-Mamba feature enhancement modules connected sequentially. See details. Figure 2The CNN-Mamba feature enhancement module consists of consecutively connected CNN modules and Mamba modules. Thus, the self-supervised reconstruction module can be viewed as a series of CNN module 1, Mamba module 1, CNN module 2, Mamba module 2, CNN module 3, Mamba module 3, CNN module 4, Mamba module 4, CNN module 5, and Mamba module 5. CNN modules 1-5 have identical structures, and so do Mamba modules 1-5.

[0040] For details on the CNN-Mamba feature augmentation module, please refer to [link / reference]. Figure 3 In the CNN-Mamba feature enhancement module, the CNN module consists of sequentially connected transposed convolutional layers, the GELU activation function, and layer normalization. The Mamba module includes a sequentially connected first linear projection layer (…). Figure 3 (Denotes as linear projection layer 1), convolutional layer, and first nonlinear transformation layer ( Figure 3 Noted as nonlinear transformation 1), state-space model SSM and third linear projection layer ( Figure 3 (Refered as linear projection layer 3); also includes a second linear projection layer ( Figure 3 (denoted as linear projection layer 2) and the second nonlinear transformation layer ( Figure 3 This is denoted as nonlinear transformation 2).

[0041] In the CNN-Mamba feature enhancement module, the output of the layer normalization is used as the output of the CNN module and connected to the first and second linear projection layers respectively. The output of the second linear projection layer is connected to the input of the second nonlinear transformation layer. The output of the second nonlinear transformation layer is multiplied element-wise with the output of the state-space model (SSM) and then used as the input of the third linear projection layer. The output of the third linear projection layer serves as the output of both the Mamba module and the CNN-Mamba feature enhancement module.

[0042] The cascaded MT architecture comprises four cascaded feature interaction modules. These modules have identical structures, consisting of consecutively connected Res-Attention and Res-Mamba modules. Specifically, the cascaded MT architecture is composed of Res-Attention module 1, Res-Mamba module 1, Res-Attention module 2, Res-Mamba module 3, Res-Mamba module 4, and Res-Mamba module 4 connected in series. Res-Attention modules 1-4 have identical structures, as do Res-Mamba modules 1-4. The input to Res-Attention module 1 is connected to the output of the self-supervised reconstruction.

[0043] The Res-Attention module includes a sequentially connected 3D Fast Fourier Transform layer, a linear projection layer, an attention calculation layer, an output projection layer, and a first normalization layer. Figure 5 The layer shown is normalized layer 1), feedforward network, and second normalized layer ( Figure 5 The layer shown is normalized (2) and the dimension reduction layer.

[0044] The attention computation layer employs multiple parallel single-head attention branches. The outputs of these branches are concatenated and projected to obtain the attention fusion features. The input features of the Res-Attention module undergo a 2D Fast Fourier Transform and linear projection, followed by attention head splitting. Each attention head processes its input, and the outputs of these heads are then concatenated and passed through an output projection layer to obtain the output projection. The output projection is processed by a first normalization layer, and then a residual is calculated between the output projection and the normalized output. This residual feature is then fed into a feedforward network (FFN) for further processing. The output of the FFN is processed by a second normalization layer, and then a residual is calculated between the output of the FFN and the normalized output. This residual feature is then dimensionality-restored to obtain the output of the Res-Attention module.

[0045] Within the same feature interaction module, the output of the Res-Attention module is connected to the input of the Res-Mamba module.

[0046] The Res-Mamba module includes a sequentially joined dimension flattening layer, a State Space Model (SSM), and a third-layer normalization layer. Figure 5 (See layer normalization 3). After the input of the Res-Mamba module is flattened, it is processed by the input state space model SSM. The output features of the state space model SSM are projected onto the input feature space of the third normalization layer. The normalized output of the third normalization layer is residuald with its input features, and then the dimensions are restored to obtain the output of the Res-Mamba module, which is also the output of the feature interaction module, so as to serve as the input of the next feature interaction module.

[0047] In the serial MT architecture, the outputs of the four feature interaction modules from shallow to deep are denoted as feature 1, feature 2, feature 3 and feature 4, respectively.

[0048] like Figure 6As shown, the frequency-aware enhancement module includes a sequentially connected multi-band feature decomposition network, a two-dimensional fast Fourier transform network, a multi-band masking separation network, and a frequency feature fusion enhancement network. The input to the frequency-aware enhancement module is the concatenated feature 1 and feature 2 output from the cascaded MT architecture. The frequency-aware enhancement module divides the input concatenated feature frequency bands into low-frequency, power frequency, and high-frequency features based on a frequency threshold and unifies the dimensions; then, it performs Fourier transforms on each to obtain low-frequency, power frequency, and high-frequency features respectively; next, it applies a frequency band mask to each frequency domain feature according to the set frequency allocation mechanism to obtain the target frequency features for each band; finally, the frequency feature fusion enhancement network uses adaptive weights to fuse the target frequency features of each band to generate enhanced features.

[0049] At the tail of the frequency-aware enhancement module, shallow and deep enhancement features can be separated. A global modeling mechanism is defined and added, and its mathematical expression is:

[0050] In the formula, Indicates the first High-frequency feature map within a local region or frequency band. Indicates the first Low-frequency characteristic map within a local region or frequency band. This indicates global average pooling, used to compress the features of the corresponding frequency band into a global description vector; For a global description of high-frequency information, For a global description of low-frequency information, Low-frequency information For high-frequency information, For the domain range.

[0051] A frequency domain allocation mechanism is defined and added between the two-dimensional fast Fourier transform and multi-band mask segmentation in the frequency sensing enhancement module. Its mathematical expression is as follows:

[0052] In the formula, This represents local curvature, i.e., the characteristics of a normal manifold. The intensity of second-order variation in the local frequency domain neighborhood is used to characterize the degree of characteristic variation in the local frequency region. For high-frequency cutoff radius, These are characteristics of the normal manifold domain.

[0053] Specifically, the local region can be determined in the frequency domain coordinates after a two-dimensional FFT. (Frequency domain coordinates) Centered on, with a radius of The neighborhood is denoted as the local region. :

[0054] Within this neighborhood, the local curvature is estimated using the second-order difference or discrete Laplace response. :

[0055] in, The response difference is calculated by the characteristic response difference between adjacent coordinates in the frequency domain neighborhood. This is determined locally by a frequency band neighborhood obtained through frequency domain masking or a sliding window. This is estimated from the second-order variation intensity of the normal manifold characteristics within this neighborhood. The above formula can be further written as:

[0056] in, , , , and These represent the normal manifold feature z in the frequency domain coordinates. , , , and The frequency domain characteristic response value at that coordinate. If z is a multi-channel characteristic, the discrete Laplace response can be calculated for each channel separately, or the channel mean / norm can be used as the comprehensive response value at that coordinate.

[0057] The text encoder processes the input text prompts, which include normal text prompts and abnormal text prompts. The text encoder processes the normal text prompts and abnormal text prompts separately, generating normal text embeddings and abnormal text embeddings respectively, which together constitute the text features.

[0058] like Figure 9 As shown, the method for constructing a defect detection model includes the following steps: S1. First, a self-supervised reconstruction module is pre-trained on a normal image dataset. The defect detection model processes normal image samples, and then the reconstruction loss is used to update the self-supervised reconstruction module. Reconstruction losses for:

[0059] in, This represents the intermediate layer features output by the serial MT architecture, namely the fused features of features 1 and features 2. The fusion method can be to first concatenate and then perform convolution processing. This represents the enhanced characteristics of the frequency-aware enhancement module output and the intermediate layer characteristics of the serial MT architecture. The reconstructed features obtained after fusion; This represents a set of normal pixel or normal feature locations. express The number of elements, Indicates the location of a pixel or feature.

[0060] S2. Train the defect detection model on a learning dataset that contains both normal and defective image samples, and update the self-supervised reconstruction module, cascaded MT architecture, frequency-aware enhancement module, and text encoder; use the final loss function for training. for:

[0061]

[0062] in, Indicates the total loss. The weighting coefficients for reconstruction loss; , and These are the weighting coefficients. Image features and text features Alignment loss, Image features and reconstruction features Alignment loss, For reconstruction features and text features Alignment loss.

[0063] This method employs a two-stage training and joint fine-tuning strategy. In the first stage, the self-supervised reconstruction module is pre-trained using only normal samples and reconstruction loss, enabling it to learn the normal sample manifold and possess high-fidelity reconstruction capabilities, thus providing more reliable reconstructed images. In the second stage, cross-modal alignment and anomaly detection capabilities are simultaneously enhanced.

[0064] The two-stage approach helps improve the convergence efficiency of model training. Introducing reconstruction loss into the final loss can further serve as an auxiliary constraint term in the joint optimization.

[0065] During the second phase of training, text prompts include normal text prompts and abnormal text prompts.

[0066] In this embodiment, the alignment loss can specifically be cosine distance. For example, the alignment loss between features A and B can be denoted as... :

[0067] If normal and abnormal text prompts are involved, a contrastive cross-entropy loss can be further constructed based on the cosine similarity mentioned above.

[0068] In this embodiment, a basic model is built according to the module connection diagram of SAFER-CLIP (Figure 1(a), Figure 1(b)); then, through self-supervised learning on a public dataset, the converged basic model is used as an anomaly detection model.

[0069] In practice, the anomaly detection model is constructed as follows: first, a basic model is built and a public dataset is downloaded; then, machine learning methods are used to enable the basic model to perform detection on the public dataset.

[0070] Datasets: Detection was performed on 13 publicly available datasets, covering various industrial detection scenarios and medical imaging fields, to analyze the detection performance of SAFER-CLIP. For industrial detection, detection was conducted on the following 7 public datasets: MVTec AD, VisA, MPDD, BTAD, SDD, DAGM, and DTD-Synthetic. For medical imaging, detection was conducted on the following 6 public datasets: CVC-ClinicDB and CVC-ColonDB for colon polyps, TN3k for thyroid nodules, BrainMRI and Br35H for brain tumors, and COVID-19 for COVID-19.

[0071] Implementation details: The publicly available CLIP visual encoder (VIT-L / 14@336px) is used, with an additional Mamba module as the backbone network. Crucially, unlike typical setups that might unfreeze the entire encoder, the parameters of the original CLIP visual encoder are specifically frozen, while only the parameters of the newly added Mamba module are unfrozen for adaptive learning. Learnable word embeddings are supported. The length is set to 12. Learnable token embeddings are appended to the first 9 layers of the text encoder to refine the text space. The value was set to 4. The top-level feature map was also used for anomaly segmentation as local visual details. The results for the statistical report datasets are the averages across their respective subsets. All experiments were performed using PyTorch-2.0.0 and a single NVIDIA A40 GPU.

[0072] Evaluation metrics: ZSAD performance was evaluated using pixel-level and image-level area under the receiver operating characteristic (AUROC). For comprehensive analysis, image-level mean precision (AP) was also included. Anomaly segmentation was evaluated using pixel-level area under the precision-recall curve (AUPRC), AAUOC, and per-region overlap (PRO).

[0073] The experimental results are shown in Table 1 below, where SAFER-CLIP represents the method of this invention, and the experimental results of the other methods are cited from the AnomalyCLIP paper and the GenCLIP paper, respectively.

[0074] Table 1: Performance Comparison of Zero-Sample Anomaly Detection on Industrial Datasets

[0075] Table 1 comprehensively presents the zero-shot anomaly detection results of current state-of-the-art models (CLIP, WinCLIP, VAND, AdaCLIP, GenCLIP, AnomalyCLIP, and SAFER-CLIP) on seven industrial datasets in complex industrial environments. AnomalyCLIP aims to learn object-independent semantic textual cues, but its performance in distinguishing between anomalies and edges remains unsatisfactory. The SAFER-CLIP model enhances anomaly sensitivity through residual reconstruction achieved via a self-supervised reconstruction mechanism based on the manifold assumption, and improves boundary and edge anomaly recognition capabilities using a frequency-domain-aware feature enhancement module, achieving the most competitive performance in the AUROC metric across datasets. On multiple datasets, SAFER-CLIP consistently outperforms other models in image-level AUROC and AP metrics. On the MVTec AD dataset, SAFER-CLIP achieves an AUROC of 93.2%, while AnomalyCLIP achieves 91.5%. The experimental results demonstrate the powerful image-level anomaly detection capabilities of SAFER-CLIP. The advanced fusion strategy proposed in this invention significantly improves the accuracy and stability of the SAFER-CLIP model for anomaly detection in more challenging industrial environments.

[0076] Figure 7 This section provides visualizations of anomaly detection. The first row displays the original anomaly image, while the second through fifth rows present images generated by SAFER-CLIP and other models.

[0077] Figure 7 The visualization results show that the SAFER-CLIP model excels in accurately locating anomalies and suppressing false positives caused by high-frequency normal textures and edges. For high-texture mesh images, AnomalyCLIP incorrectly labels a large number of normal pattern regions as anomalies. In stark contrast, the SAFER-CLIP model accurately identifies true defects while maintaining the cleanliness and normality of the background areas, demonstrating that the fusion strategy of the SAFER-CLIP model can effectively distinguish anomalies from significant high-frequency features.

[0078] Figures 8(a)-(d) show the anomaly detection results of SAFER-CLIP in four different types of images: cable, capsule, mesh, and bottle. The first row shows the input image, the second row shows the ground truth anomaly mask, where white pixels represent anomaly regions, and the third row shows the model's anomaly score map, where warmer colors indicate higher anomaly scores for detected anomaly regions.

[0079] Figures 8(a)-(d) show the experimental data, demonstrating that SAFER-CLIP can effectively detect industrial objects with varying texture complexities, including high-texture meshes and low-texture capsules. Through innovative design combining frequency-aware feature enhancement and a serial MT architecture, SAFER-CLIP successfully addresses the false alarm problem of AnomalyCLIP in high-frequency normal texture regions and effectively improves the clarity of defect localization. Furthermore, SAFER-CLIP enhances the model's sensitivity to anomalies through a self-supervised reconstruction mechanism based on the manifold assumption, thereby improving detection accuracy, especially in cases of complex textures and blurred edges.

[0080] To further evaluate the generalization ability of the SAFER-CLIP model, a comprehensive experiment was conducted on six professional medical datasets. The experimental results are detailed in Table 2.

[0081] Table 2 Comparison of zero-sample anomaly detection performance on medical datasets

[0082] The results show that SAFER-CLIP demonstrates significant improvements over AnomalyCLIP across various average metrics on six medical datasets, consistently achieving superior performance. SAFER-CLIP employs a cascaded MT architecture, which effectively captures long-range dependencies and promotes selective high-level semantic alignment. In contrast, AnomalyCLIP tends to lose local details during global feature interactions, leading to insufficient localization accuracy. Furthermore, by integrating a frequency-aware enhancement module, which utilizes multi-frequency decomposition and adaptive feature fusion techniques, SAFER-CLIP significantly enhances the model's ability to distinguish anomalies from edges. On BrainMRI, this resulted in SAFER-CLIP increasing the AUROC from 90.3% for AnomalyCLIP to 94.7%. These advancements enable the SAFER-CLIP model to adapt more robustly and accurately to medical image data influenced by various complex features, demonstrating its powerful generalization ability in recognizing complex and diverse features.

[0083] Ablation experiment of the model

[0084] The following four ablation experimental models were built: SAFER-CLIP . The .T model removes the series-connected MT architecture from the SAFER-CLIP model; SAFER-CLIP . The .S model removes the self-supervised reconstruction mechanism from SAFER-CLIP; SAFER-CLIP . The .F model removes the frequency-aware enhancement module from SAFER-CLIP; AnomalyCLIP .F adds a frequency awareness enhancement module to AnomalyCLIP.

[0085] Table 3 Performance comparison of SAFER-CLIP, AnomalyCLIP and their ablation variants

[0086] Table 3 compares AnomalyCLIP and SAFER-CLIP using MVTec AD and VisA as test sets. . T, SAFER-CLIP . Performance of four model variants: .S, SAFER-CLIP, and SAFER-CLIP. In MVTec AD, the SAFER-CLIP model achieved the best pixel-level AUROC (92.9%) and PRO (86.6%), compared to the baseline model SAFER-CLIP. . Compared to the .T model's AUROC of 91.3% and PRO of 83.9%, the model achieved a significant and stable improvement in anomaly detection metrics after fully integrating this module. Test results of the SAFER-CLIP model show that this unique fusion mechanism, through synergistic effects, not only significantly optimizes the boundary between normal data manifolds and anomalous patterns, thus greatly enhancing their separability, but also significantly improves the model's robustness to distribution shifts. Simultaneously, this fusion mechanism effectively achieves long-range dependency modeling and selective high-level semantic alignment by capturing more subtle anomalous patterns and enhancing feature representation, thereby significantly improving the model's overall anomaly detection performance and generalization ability. The self-supervised reconstruction mechanism innovatively designs a geometric manifold reconstruction module. The core mechanism of this module lies in achieving high-fidelity reconstruction of normal data on a predefined manifold, thus providing a clear basis for anomaly detection. Based on this, a dynamically adaptive anomaly threshold mechanism is further introduced, which can autonomously adjust according to the sample content, effectively improving the accuracy of anomaly detection.

[0087] The experimental results in Table 3 clearly and strongly demonstrate the effectiveness of the self-supervised reconstruction mechanism as a key module, as evidenced by its value in constructing stable and accurate representations of normal data manifolds and identifying anomalies. It has been verified that this module can directly interpret the reconstruction residuals as geometric deviations from the normal data manifold, thereby improving the model's sensitivity to anomalies. By ensuring high-quality reconstruction of normal patterns, it enhances the overall capability of zero-shot anomaly detection.

[0088] Table 4 Ablation Experiment Results of Frequency Sensing Enhancement Module

[0089] The core innovation of the frequency-aware enhancement module lies in its ingenious utilization of the energy surge characteristic of anomalous signals within a specific frequency band. To this end, the frequency-aware enhancement module designed in this invention aims to accurately capture anomalies in the frequency domain. Based on Passevar's identity, drastic local structural changes caused by anomalies manifest as a significant increase in the energy of high-frequency components. Based on this frequency domain analysis, this invention constructs an adaptive multi-level feature fusion strategy. This strategy adaptively weights and fuses features from different encoder stages in the frequency domain, focusing on specific frequency domain energy changes caused by anomalies, thereby comprehensively and accurately capturing various forms of anomaly patterns. This step is particularly crucial, as it enhances the model's ability to distinguish anomalies from edges in the high-frequency region, effectively preventing misjudgments. The ablation experiments in Table 4 clearly reveal the importance of this module and its contribution to the overall model performance. In MVTec AD, the pixel-level PRO of the complete model reaches 86.6%, an improvement from 83.8% (without the frequency-aware enhancement module); in VisA, the pixel-level PRO of the complete model increases from 86.6% to 89.3%, and the image-level AUROC increases from 82.5% to 85.3%. Data shows that traditional models such as AnomalyCLIP have a core weakness in terms of insufficient positioning accuracy in fine-textured scenes, and the frequency-aware enhancement module precisely makes up for this deficiency.

[0090] To further test the model's generalization ability in anomaly detection in industrial scenarios, experiments were conducted on the following 16 public datasets of industrial objects and materials: Carpet, Bottle, Hazelnut, Leather, Cable, Capsule, Grid, Pill, Transistor, Metal nut, Screw, Toothbrush, Zipper, Tile, Wood, and mean. The experimental results are shown in Table 5.

[0091] Table 5. Pixel-level PRO index results of the frequency sensing enhancement module based on the AnomalyCLIP framework.

[0092] Furthermore, Table 5 demonstrates the broad applicability and flexibility of the frequency-aware enhancement module; its integration into the baseline model AnomalyCLIP significantly improves pixel-level PRO performance. This highlights the clear effectiveness, universality, and scalability of frequency domain detection results. These important findings collectively validate the proposed anomalous pattern capture strategy based on frequency domain anomalous energy changes, affirming its superior effectiveness, critical importance, and broad applicability. Experimental results show that performance is improved across all categories after integrating the frequency-aware feature enhancement module, with the average PRO increasing from 81.4% to 83.9%; the Hazelnut category shows the most significant improvement, increasing from 92.4% to 95.5%, and the Wood category from 91.2% to 94.8%. The frequency-aware enhancement module is independent of any specific model architecture and can directly enhance the sensitivity of AnomalyCLIP to frequency domain features. The original architecture of AnomalyCLIP has a weak ability to distinguish high-frequency features, resulting in blurred boundaries between normal textures and abnormal defects. The frequency-aware enhancement module strengthens the difference between the two through frequency domain energy analysis. Therefore, even when migrated to the baseline model, it can still effectively improve the positioning accuracy, verifying the general enhancement value of the frequency-aware enhancement module.

[0093] To comprehensively and rigorously verify the performance advantages and generalization ability of the SAFER-CLIP model in anomaly detection tasks based on actual industrial needs, this study conducted comparative experiments on the public dataset VisA (Candle, Capsules, Cashew, Chewingg, Fryum, Macaroni1, Macaroni2, Pcb1, Pcb2, Pcb3, Pcb4, Pipe.fryum) to fully evaluate the model's detection performance. The VisA dataset contains images of various industrial objects with different texture complexities and background variations, simulating complex scenes in real-world industrial environments.

[0094] Table 6. Performance comparison of AUROC metrics for fine-grained subsets of anomaly classification on the VisA dataset.

[0095] Table 6 uses image-level AUROC as an indicator to demonstrate the zero-shot classification capabilities of the SAFER-CLIP model compared to six other models: CLIP, CLIP-AC, WinCLIP, VAND, CoOp, and AnomalyCLIP. The data shows that SAFER-CLIP achieves an average AUROC of 85.3%, a 3.2% improvement over AnomalyCLIP's 82.1%, and significantly outperforms subsets such as Capsules' 81.5% and Pcb4's 93.9%, even far exceeding the base model CLIP's 66.4%, demonstrating its accuracy advantage in fine-grained defect classification of industrial objects.

[0096] Table 7 Comparison of AP metrics for fine-grained subsets of anomaly classification on the VisA dataset.

[0097] Table 7 uses AP as an indicator to further evaluate the SAFER-CLIP model's ability to balance precision and recall in scenarios with scarce outlier samples. Experimental results show that SAFER-CLIP's average AP is 85.4%, on par with AnomalyCLIP, but it performs better on more difficult subsets: Capsules' AP is 91.9%, a 3.2% improvement over AnomalyCLIP's 88.7%, and Pipe.fryum's AP reaches 97.1%, slightly higher than AnomalyCLIP's 96.3%. Even on niche subsets like Macaroni2, SAFER-CLIP outperforms CLIP's 65%, demonstrating the SAFER-CLIP model's good generalization stability.

[0098] Table 8. Performance comparison of AUROC metrics for fine-grained subsets of anomaly segmentation on the VisA dataset.

[0099] Table 8 focuses on pixel-level segmentation tasks, using pixel-level AUROC to measure the ability to distinguish abnormal and normal regions pixel by pixel. SAFER-CLIP achieves an average pixel-level AUROC of 96.4%, a 0.9% improvement over AnomalyCLIP's 95.5% and a 16.8% improvement over WinCLIP's 79.6%. It maintains high accuracy even in structurally complex Pcb subsets. In particular, in the Capsules subset, SAFER-CLIP accurately captures the pixel-level location of subtle surface defects, validating its ability to resist interference from complex textures of industrial objects.

[0100] Table 9. Performance comparison of fine-grained subset PRO metrics for anomaly segmentation on the VisA dataset.

[0101] Table 9 focuses on PRO (Probability of Error) to evaluate the accuracy of anomaly area localization. SAFER-CLIP achieved an average PRO of 96.4%, far exceeding AnomalyCLIP's 87% and CLIP's 14.8%. The Candle subset PRO reached 98.4%, a 2.2% improvement over AnomalyCLIP's 96.2%. Furthermore, SAFER-CLIP achieved a high PRO of 98.2% on the Pipe.fryum subset, accurately locating pipe surface defects. Even on low-performance subsets such as Pcb2, it outperformed CLIP's 12.8%, fully demonstrating its robustness in pixel-level anomaly localization in industrial scenarios.

[0102] Through comprehensive comparative experiments on the VisA dataset, the anomaly detection capabilities of SAFER-CLIP were fully validated. Experimental results show that SAFER-CLIP demonstrates significant advantages over other existing methods in various industrial anomaly detection tasks, particularly in anomaly localization accuracy against high-texture and complex backgrounds, showcasing its excellent generalization ability. The experiments also demonstrate the potential of the SAFER-CLIP model for practical applications in the industrial field, especially its adaptability to complex object shapes, texture complexity, and diverse backgrounds, providing an effective solution for industrial anomaly detection tasks.

Claims

1. A method for constructing a zero-sample defect detection model in the industrial field, characterized in that, First, a visual language model is constructed, comprising: a self-supervised reconstruction module, a frequency-aware enhancement module, a cascaded MT architecture, a text encoder, and a total loss module. The self-supervised reconstruction module reconstructs the input image. The reconstructed image is processed by the cascaded MT architecture, which outputs multiple layers of features from shallow to deep. The shallowest and deepest features are processed by the frequency-aware enhancement module to obtain enhanced features. The enhanced features output by the frequency-aware enhancement module are fused with the intermediate layer features output by the cascaded MT architecture to obtain the reconstructed features. The deepest feature output by the cascaded MT architecture is also output as an image feature. The text encoder processes the set text prompts to generate text features that tend towards the image category. The text prompts include normal text prompts and abnormal text prompts, and the image categories include normal image samples and defective image samples. Normal images are defect-free image samples. The total loss module aligns the reconstructed features, image features, and text features pairwise, calculates the alignment loss, and fuses the three alignment losses to obtain the total loss function. During training, normal images are extracted as training samples and input into the visual language model. After processing, the total loss function is obtained. The visual language model is then updated based on the direction of the total loss function. Repeat the above steps until the visual language model converges.

2. The method for constructing a zero-sample defect detection model in the industrial field as described in claim 1, characterized in that, The defect detection model first trains a self-supervised reconstruction module on a normal image dataset; then it trains the defect detection model on a learning dataset that contains both normal image samples and defect image samples, and updates the self-supervised reconstruction module, the cascaded MT architecture, the frequency-aware enhancement module, and the text encoder. During the training of the self-supervised reconstruction module, the loss function adopts the alignment loss between the intermediate layer features output by the cascaded MT architecture and the reconstructed features; The loss function used in the overall model training process includes image features. Reconstruction features and text features The total loss is obtained by weighted summation of the alignment losses between each pair of elements.

3. The method for constructing a zero-sample defect detection model in the industrial field as described in claim 2, characterized in that, Loss function used during overall model training for: in, Indicates the total loss. To rebuild the losses Weighting coefficients, reconstruction loss Alignment loss between reconstructed features and intermediate layer features output by the serial MT architecture.

4. The method for constructing a zero-sample defect detection model in the industrial field as described in claim 1, characterized in that, The self-supervised reconstruction module consists of five identical CNN-Mamba feature enhancement modules connected sequentially. The CNN-Mamba feature enhancement module consists of CNN modules and Mamba modules connected sequentially. The CNN module consists of sequentially connected transposed convolutional layers, the GELU activation function, and layer normalization; the Mamba module includes a sequentially connected first linear projection layer, convolutional layers, a first nonlinear transformation layer, a state-space model SSM, and a third linear projection layer; it also includes a second linear projection layer and a second nonlinear transformation layer; The output of the normalized layer is connected to the first linear projection layer and the second linear projection layer as the output of the CNN module. The output of the second linear projection layer is connected to the input of the second nonlinear transformation layer. The output of the second nonlinear transformation layer is multiplied element-wise with the output of the state-space model SSM and then used as the input of the third linear projection layer. The output of the third linear projection layer is used as the output of the CNN-Mamba feature enhancement module.

5. The method for constructing a zero-sample defect detection model in the industrial field as described in claim 1, characterized in that, The serial MT architecture consists of four identical feature interaction modules connected in series. The feature interaction module consists of a Res-Attention module and a Res-Mamba module connected sequentially. The Res-Attention module includes a sequentially connected 3D Fast Fourier Transform layer, a linear projection layer, an attention computation layer, an output projection layer, a first normalization layer, a feedforward network, a second normalization layer, and a dimension reduction layer. The attention computation layer employs multiple parallel single-head attention branches; The input features of the Res-Attention module are processed by a 2D Fast Fourier Transform and linear projection, then split into attention heads. Each attention head processes its input, and the outputs of each attention head are concatenated and passed through an output projection layer to obtain the output projection. The output projection is processed by a first normalization layer and then a residual is calculated with the output projection. The residual features are fed into a feedforward network (FFN) for processing. The output of the FFN is processed by a second normalization layer and then a residual is calculated with the output of the FFN. The dimensionality of the residual features is then restored to obtain the output of the Res-Attention module. The Res-Mamba module consists of a sequentially connected dimension-flattening layer, a state-space model (SSM), and a third normalization layer. After dimension flattening, the input of the Res-Mamba module is processed by the state-space model (SSM). The output features of the state-space model (SSM) are projected onto the input feature space of the third normalization layer. The normalized output of the third normalization layer is left as a residual with its input features, and then the dimensions are restored to obtain the output of the Res-Mamba module.

6. The method for constructing a zero-sample defect detection model in the industrial field as described in claim 4, characterized in that, The frequency awareness enhancement module divides the input spliced ​​feature frequency bands into low-frequency features, power frequency features, and high-frequency features according to the frequency threshold and unifies the dimensions; then it performs Fourier transform on each feature to obtain low-frequency frequency domain features, power frequency frequency domain features, and high-frequency frequency domain features respectively; and then applies a frequency band mask to each frequency domain feature according to the set frequency domain allocation mechanism to obtain the target frequency features of each frequency band. Then, the frequency feature fusion enhancement network uses adaptive weights to fuse the target frequency features of each frequency band to generate enhanced features.

7. The method for constructing a zero-sample defect detection model in the industrial field as described in any one of claims 1-6, characterized in that, The defect detection model consists of a trained visual language model and a detection output module. The detection output module is used to calculate a comprehensive anomaly score by weighted summation of the distances between each pair of reconstructed features, image features, and text features. The image samples to be detected are judged based on the comprehensive anomaly score. Is it a normal image or a defective image? Comprehensive Abnormal Score For reconstruction features Image features and text features Weighted summation of the distances between each pair of pairs.

8. The method for constructing a zero-sample defect detection model in the industrial field as described in claim 7, characterized in that, The detection output module will combine the anomaly scores. With threshold Compare; If the overall anomaly score is less than the threshold Then determine the input image sample to be detected. Normal; otherwise, it is judged as abnormal, that is, there is a defect. threshold Take the 95th percentile of the comprehensive abnormal scores of normal samples.

9. The method for constructing a zero-sample defect detection model in the industrial field as described in claim 8, characterized in that, threshold , and These are the mean and standard deviation of the combined outlier scores on the normal validation set, respectively. This is an empirical coefficient.

10. A system for constructing a zero-sample defect detection model in the industrial field, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, the processor is connected to the memory, and the processor is used to execute the computer program to implement the method for constructing an industrial zero-sample defect detection model as described in any one of claims 1-7.