A Few-Shot Industrial Anomaly Detection Method Based on Cross-Modal Adaptive Interaction

Through the cross-modal adaptive interaction method, combined with the field knowledge injection of dual-path vision encoder and large language model, the subtle anomaly recognition problem in complex backgrounds in industrial anomaly detection is solved, and high-precision and low-cost industrial anomaly detection are achieved.

CN119989247BActive Publication Date: 2025-07-25NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510484103.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-25
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing industrial anomaly detection methods are difficult to distinguish subtle anomalies under complex backgrounds, are inefficient in cross-modal interactions, and have weak semantic adaptability of static text templates, resulting in high missed detection rates and it is difficult to achieve high-precision detection under the condition of few samples.

Method used

Using a method based on cross-modal adaptive interaction, a dual-path vision encoder and a multi-level feature adaptive fusion adapter, combined with a cross-modal dynamic prompt embedded device and interaction module, the domain knowledge injection of large language models is used to achieve efficient alignment of visual and text features and accurate recognition of abnormal features.

Benefits of technology

In the case where only a small number of normal samples are required, the detection sensitivity and accuracy of micro defects is significantly improved, the leakage detection rate is reduced, and a low-cost and high-precision industrial anomaly detection solution is provided, supporting image-level and pixel-level abnormality detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989247B_ABST
    Figure CN119989247B_ABST
Patent Text Reader

Abstract

The present invention discloses a few-shot industrial anomaly detection method based on cross-modal adaptive interaction. According to the aligned final features, normal semantic projection, abnormal semantic projection, optimized visual features, and optimized text features, the total loss is calculated, and the total loss is used to update the parameters of the industrial anomaly detection model, obtaining a trained industrial anomaly detection model; the industrial image to be detected is input into the trained industrial anomaly detection model to obtain the aligned final features, abnormal semantic projection, optimized visual features, and optimized text features, which are used to judge the anomaly situation of the industrial image to be detected. The present invention can significantly enhance the model's ability to distinguish normal and abnormal features under the condition of only requiring a small number of normal samples, reduce the dependence on labeled data, and provide a high-precision, low-cost, and quickly deployable industrial anomaly detection solution for intelligent manufacturing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a few-shot industrial anomaly detection method based on cross-modal adaptive interaction, belonging to the technical field of industrial detection. Background Art

[0002] Industrial anomaly detection is a core link in intelligent manufacturing, aiming to identify and locate defects or abnormal situations in products or processes through automated means. Traditional methods mostly rely on supervised learning and require a large number of labeled anomaly samples. However, anomaly samples are scarce in industrial scenarios and the labeling cost is high, resulting in the model being difficult to adapt to the actual industrial scenario.

[0003] In recent years, few-shot anomaly detection (FSAD) has become a research direction that has received much attention. FSAD aims to train a model using limited normal samples and infer the existence of anomalies by learning the distribution characteristics of normal data. Compared with traditional methods, few-shot learning can achieve higher detection performance in data-limited scenarios, so it has important industrial application potential.

[0004] However, the implementation of few-shot anomaly detection still faces technical challenges, especially in how to efficiently extract anomaly features and improve detection accuracy. At the same time, multi-modal vision-language models (Vision-Language Models, VLM, such as CLIP: Contrastive Language-Image Pre-Training) have made remarkable progress in the cross field of computer vision and natural language processing. These models show powerful performance in tasks such as image classification and object detection by fusing visual information (such as images) and text information (such as descriptive language). The core advantage of VLM lies in its cross-modal ability, that is, guiding visual feature extraction through text or enhancing semantic understanding in reverse through visual information. Therefore, applying VLM with cross-modal ability to industrial anomaly detection, especially in few-shot scenarios, has significant practical value.

[0005] However, there are some defects in existing anomaly detection algorithms:

[0006] (1) Single-modal limitation: Existing traditional methods are mostly based on pure visual features (such as image segmentation, reconstruction error), lacking in-depth understanding of defect semantics and being difficult to distinguish subtle anomalies in complex backgrounds.

[0007] (2) Inefficient cross-modal interaction: A few multi-modal methods that combine vision and text have problems such as rough feature fusion and inaccurate modal alignment, resulting in insufficient local defect response. How to effectively guide the interaction and mutual prompting between text information and visual features and achieve accurate anomaly recognition is still an urgent problem to be solved.

[0008] (3) Insufficient semantic information: Most existing multi-modal methods use static text prompt templates (such as "a photo of [state]"), lacking refined domain knowledge injection and being difficult to dynamically adapt to the semantic requirements of different industrial scenarios, resulting in an increased missed detection rate. Summary of the Invention

[0009] Objective: To overcome the problems of complex background interference caused by the limitations of single-modal features, insufficient local defect responses caused by inefficient cross-modal interactions, and high missed detection rates caused by weak semantic adaptation capabilities of static text templates in the prior art, the present invention provides a few-shot industrial anomaly detection method based on cross-modal adaptive interaction.

[0010] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0011] In a first aspect, a few-shot industrial anomaly detection method based on cross-modal adaptive interaction specifically includes:

[0012] Obtain a few-shot dataset.

[0013] Input the few-shot dataset into a dual-path visual encoder to obtain dual-path residual fusion features of multiple layers.

[0014] Input the dual-path residual fusion features of multiple layers in the dual-path residual fusion features of multiple layers into a multi-level feature adaptive fusion adapter to obtain the aligned final features.

[0015] Input the aligned final features into a cross-modal dynamic prompt embedder to obtain normal semantic projections and abnormal semantic projections.

[0016] Input the aligned final features, normal semantic projections, and abnormal semantic projections into a cross-modal interaction module to obtain optimized text features and optimized visual features.

[0017] Calculate the total loss based on the aligned final features, normal semantic projections, abnormal semantic projections, optimized visual features, and optimized text features, and update the parameters of the industrial anomaly detection model using the total loss to obtain a trained industrial anomaly detection model.

[0018] Input the industrial image to be detected into the trained industrial anomaly detection model to obtain the aligned final features, abnormal semantic projections, optimized visual features, and optimized text features for judging the anomaly situation of the industrial image to be detected.

[0019] The abnormal situation is located according to the aligned final features of the industrial image to be detected, the normal semantic projection and the abnormal semantic projection, the optimized visual features and the optimized text features.

[0020] In a second aspect, a computer-readable storage medium stores a computer program, which, when executed by a processor, implements a few-sample industrial anomaly detection method based on cross-modal adaptive interaction as described in the first aspect.

[0021] According to a third aspect, a computer device includes:

[0022] Memory, used to store instructions.

[0023] The processor is used to execute the instructions so that the computer device performs the operations of the few-sample industrial anomaly detection method based on cross-modal adaptive interaction as described in the first aspect.

[0024] Beneficial effects: The invention provides a method for industrial anomaly detection based on cross-modal adaptive interaction with a small number of samples. By introducing a cross-modal bidirectional adaptive interaction mechanism and a small number of sample optimization strategy, the invention shows significant beneficial effects in the field of industrial anomaly detection. Specifically, the method includes the following:

[0025] 1. Based on the dual-path attention visual encoder and the multi-level feature adaptive fusion adapter, the model can simultaneously capture the global semantic associations and local subtle abnormal features of the image, effectively overcome complex background interference, and improve the detection sensitivity of tiny defects.

[0026] 2. By embedding dynamic text prompts and injecting domain knowledge of large language models, the problem of insufficient semantic adaptation of traditional static templates is solved, and refined, scenario-adaptive exception description templates are generated, which significantly reduces the missed detection rate and enhances the accuracy of cross-modal interaction.

[0027] 3. The bidirectional cross-modal interactive optimizer combines the multi-objective joint optimization loss function to achieve efficient alignment of visual and text features and enhancement of abnormal features. Under the condition of only 4 normal reference samples, the model can still maintain high robustness, reduce dependence on labeled data, and significantly reduce industrial quality inspection costs.

[0028] 4. The multi-scale anomaly scoring and positioning module can not only provide image-level anomaly detection and judgment, but also locate pixel-level anomalies by fusing global and local feature responses, supporting high-precision defect detection and real-time decision-making, providing a low-cost, high-precision and fast-deployment solution for smart manufacturing scenarios, and promoting the actual industrial application of few-sample learning and cross-modal technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1Schematic diagram of the process of a few-shot industrial anomaly detection method based on cross-modal adaptive interaction according to the present invention.

[0030] Figure 2 Schematic diagram of the structure of the industrial anomaly detection model of the present invention. Specific embodiments

[0031] The following describes clearly and completely the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0032] The following further illustrates the present invention with specific embodiments.

[0033] Embodiment 1:

[0034] This embodiment introduces a few-shot industrial anomaly detection method based on cross-modal adaptive interaction, as Figure 1 shown, specifically including:

[0035] Step 1: Obtain a few-shot dataset.

[0036] Further, the few-shot dataset specifically includes:

[0037] Select K normal samples of each class from the authoritative benchmark dataset for industrial anomaly detection.

[0038] Among them, the authoritative benchmark dataset for industrial anomaly detection uses MVTec AD, covering high-quality images with a resolution of 1024×1024 of 15 types of industrial products.

[0039] Further, in one embodiment, randomly select K normal samples (scaled uniformly to 512×512 in size, with RGB three channels) from each class of dataset to form a few-shot training set , , as the randomly selected reference images. The characteristics of such data are high resolution, low sample volume, and missing annotations, meeting the actual needs of scarce anomaly samples in the industrial quality inspection scenario. The output set of few-shot reference images with unified scale , as the input for subsequent feature extraction.

[0040] Step 2: Input the few-shot dataset into the dual-path visual encoder to obtain dual-path residual fusion features of the layer.

[0041] Further, step 2 specifically includes:

[0042] Step 2.1: Perform convolution block by block in sequence to obtain block features , and the expression is as follows:

[0043]

[0044] Where: is a 16×16 convolution kernel with a stride of 16 to divide the input image into blocks (each block has a resolution of 16×16), and i is the sample extracted for each class, with a range of 1 - K.

[0045] Step 2.2: Perform layer normalization on the block features in sequence along the channel dimension to obtain the initial block embedding features , and the expression is as follows:

[0046]

[0047] Step 2.3: Input the initial block embedding features into the global context attention path, and generate a triple vector , and for the first - layer attention calculation through three groups of independent linear projections. The expression is as follows:

[0048]

[0049] Where: is the trainable query vector parameter, is the trainable key vector parameter, is the trainable value vector parameter.

[0050] Step 2.4: Calculate the first - layer attention features , and the expression is as follows:

[0051]

[0052] Where, is the normalization operation, is the scaling factor to prevent the inner product from being too large and causing unstable gradients, and T represents the transpose matrix.

[0053] Step 2.5: Input the first - layer attention features into the feed - forward network to expand the feature dimension and obtain the intermediate - layer features , the first - layer QKV attention , and the expression is as follows:

[0054]

[0055]

[0056] Among them, is the weight parameter of the feedforward network expansion layer, is the bias parameter of the feedforward network expansion layer, is the weight parameter of the feedforward network compression layer, is the bias parameter of the feedforward network compression layer.

[0057] Step 2.6: Input the initial block embedding feature into the spatial-channel hybrid attention path ( ), and generate the first-layer spatial weight map through lightweight single-channel convolution , and the expression is as follows:

[0058]

[0059] Among them: is 's convolution kernel, and the number of output channels is 1, is the activation function.

[0060] Step 2.7: Calculate the first-layer global average pooling value of the initial block embedding feature , and the expression is as follows:

[0061]

[0062] Step 2.8: Project the first-layer global average pooling value to a higher dimension through a linear layer and pass it through an activation function to introduce a non-linear intermediate quantity to fit the complex channel relationship, and obtain the intermediate feature

[0063]

[0064] Among them: is the parameter for dimensionality increase of the linear transformation, is the bias parameter of the linear transformation.

[0065] Step 2.9: Project the intermediate feature to a lower dimension through a linear projection layer for recovery and pass it through an activation function to obtain the first-layer channel weight , and the expression is as follows:

[0066]

[0067] Among them: is the parameter of the dimensionality reduction matrix of the projection layer, is the bias parameter of the dimensionality reduction of the projection layer.

[0068] Step 2.10: The first-layer spatial weight map , the first-layer channel weight and the initial block embedding feature are used to output the first-layer fused feature through hybrid gating weighting, and the expression is as follows:

[0069]

[0070] where, is the gated Hadamard product to achieve regional gating attention focusing, is the element-wise multiplication to achieve channel dimension scaling.

[0071] Step 2.11: Calculate the first-layer dual-path residual fusion feature according to the first-layer QKV attention , the first-layer fused feature and the initial block embedding feature , and the expression is as follows:

[0072]

[0073] Step 2.12: Repeat steps 2.4 - 2.11 to calculate the dual-path residual fusion features of the th layers respectively, and the expression is as follows:

[0074]

[0075] Step 3: Input the dual-path residual fusion features of multiple layers in the th-layer dual-path residual fusion feature into the multi-level feature adaptive fusion adapter to obtain the aligned final feature.

[0076] Furthermore, step 3 specifically includes:

[0077] Step 3.1: Extract the dual-path residual fusion features of the th, , , , th layers from the as the hierarchical features , where . , , .

[0078] Step 3.2: Perform global average pooling on the hierarchical features respectively to obtain the average pooling values , where and and .

[0079]

[0080] Where: represents traversing the channels at each spatial position (i, j).

[0081] Step 3.3: Obtain the intermediate feature from the average pooling value through a lightweight mapping layer, and the expression is as follows:

[0082]

[0083] Where is the matrix parameter for dimensionality increase, expanding the channel interaction capacity, is the bias vector, and GeLU is a non-linear activation function, introducing smooth non-linearity.

[0084] Step 3.4: Calculate the hierarchical attention weights for each layer of intermediate feature , and the expression is as follows:

[0085]

[0086] Where: is the hierarchical attention weight matrix, is the hierarchical attention bias.

[0087] Step 3.5: Normalize the hierarchical attention weights to obtain the normalized hierarchical weights .

[0088]

[0089] Where: represents the exponential function with the natural constant as the base.

[0090] Step 3.6: Weightedly fuse the hierarchical feature with the normalized attention weights to obtain the fused feature .

[0091]

[0092] Step 3.7: Align the dimensionality of the fused feature through linear projection to obtain the final aligned feature .

[0093]

[0094] Among them: is a learnable alignment matrix.

[0095] Step 4: Input the finally aligned features into the cross-modal dynamic prompt embedder to obtain the normal semantic projection and the abnormal semantic projection.

[0096] Furthermore, Step 4 specifically includes:

[0097] Step 4.1: The finally aligned features obtain compressed features through convolution operations , and the expression is as follows:

[0098]

[0099] Step 4.2: Generate spatial attention aggregation weights for the compressed features through convolution and normalization layers , and the expression is as follows:

[0100]

[0101] Step 4.3: Aggregate the context of the compressed features and the spatial attention aggregation weights to obtain a global description vector , and the expression is as follows:

[0102]

[0103] Among them, is the spatial position, , respectively represent the compressed features at each spatial position and the spatial attention aggregation weights .

[0104] Step 4.4: Dimension up the global description vector through a fully connected layer and obtain the hidden layer features through an activation function , and the expression is as follows:

[0105]

[0106] Among them: is the dimension-up matrix, is the bias matrix.

[0107] Step 4.5: Adopt a split-head projection strategy to decouple the hidden layer features into S groups of independent subspaces, and each group generates a hidden vector through an independent linear layer , where, , the expression is as follows:

[0108]

[0109] Among them, is the head projection weight of the i-th component, is the head projection bias of the i-th component, is the linear transformation parameter.

[0110] Step 4.6: Concatenate the S group of hidden vectors into a dynamic prompt prefix matrix, and obtain the dynamic prompt through layer normalization:

[0111]

[0112] Step 4.7: Call the large language model to obtain industrial semantic enhanced knowledge , the expression is as follows:

[0113]

[0114] Among them: is the large language model, is to select the Z defect descriptions with the highest confidence from the output of the large language model, represents the input class identifier.

[0115] Step 4.8: Concatenate the dynamic prefix , class name and the semantics of "flawless" to obtain the normal prompt template .

[0116]

[0117] Step 4.9: Concatenate the dynamic prefix , class name and industrial semantic enhanced knowledge to obtain the abnormal prompt template .

[0118]

[0119] Among them is the statement concatenation operation, concatenating the cross-modal dynamic prompt prefix, class name and semantic enhanced knowledge.

[0120] Step 4.10: Send the normal prompt template and the abnormal prompt template into the text encoder to obtain the normal semantic projection and the abnormal semantic projection .

[0121]

[0122] Step 5: Input the aligned final features, normal semantic projection, and abnormal semantic projection into the cross-modal interaction module to obtain optimized text features and optimized visual features.

[0123] Further, Step 5 specifically includes:

[0124] Step 5.1: Let serve as positive and negative text features, and the aligned final features , positive and negative text features are respectively subjected to cross-modal semantic alignment through a normalization layer and a linear layer to obtain a triple for attention calculation , and , including the visual query attention vector for attention calculation, the normal text key attention vector , the abnormal text key attention vector , the normal text value attention vector and the abnormal text value attention vector , and the expressions are as follows:

[0125]

[0126] where: is the learnable visual query attention weight, is the normal text key attention weight, is the abnormal text key attention weight, is the normal text value attention weight, is the abnormal text value attention weight, is the normalization operation to stabilize the feature distribution.

[0127] Step 5.2: Calculate the adaptive positive attention weight and the adaptive negative attention weight for each spatial position of the visual query attention vector , and the expressions are as follows:

[0128]

[0129] where: is the adaptive positive attention weight, is the adaptive negative attention weight, is the spatial position, is the scaling factor.

[0130] Step 5.3: According to the adaptive positive attention weight , the adaptive negative attention weight , calculate the normal text features after attention interaction optimization , the abnormal text features after attention interaction optimization , and the expression is as follows:

[0131]

[0132]

[0133] Among them, are the normal text features and abnormal text features after attention interaction optimization respectively, which optimize the text features to dynamically focus on the potential abnormal regions in the image.

[0134] Step 5.4: Obtain the non-linear high-order semantic features of the abnormal text features after attention interaction optimization through the non-linear mapping layer .

[0135]

[0136]

[0137] Wherein: is the weight of the high-order semantic extension layer, is the bias of the high-order semantic extension layer. is the weight of the high-order semantic compression layer, the bias of the high-order semantic compression layer.

[0138] Step 5.5: Project the defect description generated by the industrial semantic enhancement knowledge onto the semantic space through the text encoder to generate the domain knowledge bias term , and the expression is as follows:

[0139]

[0140] Step 5.6: After adding the non-linear high-order semantic features and the domain knowledge bias term , normalize to obtain the semantically complete abnormal text features , and the expression is as follows:

[0141]

[0142] Step 5.7: Perform convolution transformation on the aligned final features to generate the visual query vector ' aligned with the text modality, and the expression is as follows:

[0143]

[0144] Step 5.8: Copy and expand the optimized text features along the spatial dimension to generate a key vector aligned with the visual feature space .

[0145]

[0146] Among them, represents the optimized text features, .

[0147] Step 5.9: Generate spatial gating weights for the visual query vector aligned with the text modality and the key vector aligned with the visual feature space through convolution and non-linear activation functions. The expression is as follows:

[0148]

[0149] Step 5.10: Project separately into H heads to obtain multi-head text value vectors . Project , to obtain multi-head visual query vectors , multi-head text key vectors . The expression is as follows:

[0150]

[0151]

[0152]

[0153] Among them: is the query projection matrix for each head, is the key projection matrix for each head, is the value projection matrix for each head.

[0154] Step 5.11: Input , , into the H-head attention for fusion to obtain a multi-head fusion vector . The expression is as follows:

[0155]

[0156]

[0157] Among them, is the multi-head output fusion weight.

[0158] Step 5.12: The spatial gating weights and the multi-head fusion vector as well as the aligned final features perform residual weighted fusion to obtain the optimized visual features :

[0159]

[0160] Among them, is the gated product.

[0161] Step 6: According to the aligned final features, normal semantic projection, abnormal semantic projection, optimized visual features, and optimized text features, calculate the total loss, and use the total loss to update the parameters of the industrial anomaly detection model to obtain the trained industrial anomaly detection model. Among them, the industrial anomaly detection model includes: a two-path visual encoder, a multi-level feature adaptive fusion adapter, a cross-modal dynamic prompt embedder, and a cross-modal interaction module connected in sequence.

[0162] Furthermore, the specific steps of Step 6 include:

[0163] Step 6.1: Initialize the optimizer, set the learning rate, weight decay, and the number of iterations, and perform forward propagation in Steps 2 - 5 to generate the aligned final features , normal semantic projection and abnormal semantic projection , optimized visual features and optimized text features .

[0164] Step 6.2: Calculate the cross-modal contrast loss according to the aligned final features , normal semantic projection and abnormal semantic projection , and the expression is as follows:

[0165]

[0166] Among them, is the cosine similarity, is the temperature coefficient, is the number of normal samples, is the batch, , , respectively represent the aligned final features, normal semantic projection, and abnormal semantic projection of the i-th normal sample, and e is the natural constant.

[0167] Step 6.3: According to the optimized visual features , optimized text features Calculate the feature distribution alignment loss function , the expression is as follows:

[0168]

[0169] in: is a set of randomly sampled spatial locations. Align the projection head for vision Align the drop shadow header for the text, Represents the feature vector of the optimized visual feature at the spatial position (h, w). Represents the norm of 2.

[0170] Step 6.4: Projection based on normal semantics Projection with exception semantics Compute explicit text boundary constraint loss , the expression is as follows:

[0171]

[0172] in: is the minimum interval threshold, Represents the maximum value function.

[0173] Step 6.5: Based on cross-modal contrast loss , feature distribution alignment loss function and explicit text boundary constraint loss Calculate total loss , the expression is as follows:

[0174]

[0175] in: is the boundary loss weight, is the alignment loss weight.

[0176] Step 6.6: Back-propagation updates the parameters of the industrial anomaly detection model, completes the number of training iterations in turn, saves the parameters of the trained dual-path visual encoder, multi-level feature adaptive fusion adapter, cross-modal dynamic prompt embedder and cross-modal interaction module, and obtains the trained industrial anomaly detection model.

[0177] Step 7: Input the industrial image to be detected into the trained industrial anomaly detection model to obtain the aligned final features, anomaly semantic projection, optimized visual features, and optimized text features, which are used to determine the anomaly of the industrial image to be detected.

[0178] Furthermore, the step 7 specifically includes:

[0179] Step 7.1: Input the industrial image to be detected into the trained industrial anomaly detection model to obtain the final aligned features , abnormal semantic projection , optimized visual features and optimized text features .

[0180] Step 7.2: Expand the optimized visual features along the spatial dimension to obtain spatial features , and the expression is as follows:

[0181]

[0182] Step 7.3: Calculate the cosine similarity between the spatial features and the abnormal semantic projection and perform spatial reshaping to obtain the initial abnormal response map :

[0183]

[0184]

[0185] Step 7.4: Upsample the initial abnormal response map and obtain the dynamically learnable weight parameter through average pooling and layer normalization, and the expression is as follows:

[0186]

[0187] Step 7.5: Perform weighted fusion on the initial abnormal response map and the weight parameter to obtain the final abnormal response map , and the expression is as follows:

[0188]

[0189] Step 7.6: Calculate the image-level anomaly score according to the aligned final features and the optimized text features , and the expression is as follows:

[0190]

[0191] where: represents the maximum pooling value parameter of the global similarity and the local similarity, is the cosine similarity.

[0192] Step 7.7: When the image-level anomaly score When it is greater than the threshold, it is determined as an abnormal situation.

[0193] Step 8: Locate the abnormal situation based on the aligned final features of the industrial image to be detected, the normal semantic projection and the abnormal semantic projection, the optimized visual features, and the optimized text features.

[0194] Further, step 8 specifically includes:

[0195] Step 8.1: Unfold the optimized visual features along the spatial dimension to obtain spatial features , and the expression is as follows:

[0196]

[0197] Step 8.2: Calculate the cosine similarity between the spatial features and the abnormal semantic projection and perform spatial reshaping to obtain the initial abnormal response map :

[0198]

[0199]

[0200] Step 8.3: Upsample the initial abnormal response map and obtain the dynamically learnable weight parameter through average pooling and layer normalization. The expression is as follows:

[0201]

[0202] Step 8.4: Perform weighted fusion on the initial abnormal response map and the weight parameter to obtain the final abnormal response map , and the expression is as follows:

[0203]

[0204] Step 8.5: Calculate the abnormal mask based on the final abnormal response map to obtain the location of the abnormal situation.

[0205]

[0206] Among them, is the threshold, and 1 is the pixel coordinate of the abnormal situation.

[0207] Example 2:

[0208] This embodiment introduces an embodiment of a few-shot industrial anomaly detection method based on cross-modal adaptive interaction, specifically including:

[0209] Step 1.1: Obtain the publicly available industrial anomaly detection dataset MVTec AD, which contains high-quality color images of 15 industrial detection object categories, with a total of 5354 samples.

[0210] Step 1.2: Randomly select =4 samples from each category in turn to form a few-shot normal image sample set as the training set. , is the randomly selected reference image. Among them, the training set only contains defect-free normal samples.

[0211] Step 2.1: Perform convolutional block division on the output of Step 1 in turn to obtain block features , and the expression is as follows:

[0212]

[0213] Among them: is a 16×16 convolutional kernel with a stride of 16 to divide the input image into blocks (each block has a resolution of 16×16), and i is the sample extracted from each class, ranging from 1 to 4.

[0214] Step 2.2: Perform layer normalization on the block features in turn along the channel dimension to obtain the initial block embedding features , and the expression is as follows:

[0215]

[0216] Step 2.3: Input the initial block embedding features into the global context attention path, and generate a triple vector for the first-layer attention calculation through three groups of independent linear projections:

[0217]

[0218] Among them: is the trainable query vector parameter, is the trainable key vector parameter, is the trainable value vector parameter.

[0219] Step 2.4: Calculate the first-layer attention features , and the expression is as follows:

[0220]

[0221] Among them, is the normalization operation, is the scaling factor to prevent the inner product from being too large and causing unstable gradients. T represents the transposed matrix.

[0222] Step 2.5: Input the first-layer attention features into the feed-forward network to expand the feature dimension and obtain the intermediate-layer features , the first-layer QKV attention :

[0223]

[0224]

[0225] Among them, are the weight parameters of the feed-forward network expansion layer, are the bias parameters of the feed-forward network expansion layer, are the weight parameters of the feed-forward network compression layer, are the bias parameters of the feed-forward network compression layer to enhance the non-linear expression ability of the features.

[0226] is the non-linear activation function:

[0227]

[0228] Step 2.6: At the same time, input the initial block embedding features into the spatial-channel hybrid attention path ( ), and generate the first-layer spatial weight map through lightweight single-channel convolution. The expression is as follows:

[0229]

[0230] Among them: is 's convolution kernel, and the number of output channels is 1. is the activation function.

[0231] Step 2.7: Calculate the first-layer global average pooling value of the initial block embedding features , and compress each feature map into a scalar to compress the spatial dimension and generate a channel description vector.

[0232]

[0233] Step 2.8: The first-layer global average pooling value Dimensionality is increased through a linear layer for projection and non-linear intermediate quantities are introduced through the GeLU activation function to fit complex channel relationships, obtaining intermediate features

[0234]

[0235] Where: is the parameter for dimensionality increase of linear transformation, is the bias parameter of linear transformation.

[0236] Step 2.9: The intermediate features are restored by dimensionality reduction through a linear projection layer and the first-layer channel weights are obtained through the .

[0237]

[0238] Where: is the parameter of the dimensionality reduction matrix of the projection layer, is the bias parameter of dimensionality reduction of the projection layer.

[0239] Step 2.10: The first-layer spatial weight map , the first-layer channel weights and the initial block embedding features are used to output the first-layer fused features through hybrid gating weighted.

[0240]

[0241] Where, is the gated Hadamard product to achieve regional gating attention focus, is the element-wise multiplication to achieve channel dimension scaling.

[0242] Step 2.11: Calculate the first-layer dual-path residual fusion features :

[0243]

[0244] Step 2.12: Repeat steps 2.4 - 2.10 to calculate the layer , , and perform feature fusion to obtain the layer dual-path residual fusion: . In this embodiment, 12 layers are adopted, and the attention mechanism in each layer gradually extracts multi-granularity features from local details (shallow layer) to global semantics (deep layer). The 12 layers can cover the scale range of common defects in industrial anomaly detection (such as micro-cracks to structural deformations). Compared with deeper networks (such as 24 layers), the 12 layers are more suitable for the real-time requirements of industrial scenarios in terms of GPU memory occupancy and inference speed.

[0245] Step 3.1: Extract features from the 4th, 8th, and 12th layer encoding blocks of : , . Select three layers: shallow local features (local texture anomalies), middle semantic features (component-level structural anomalies), and deep global features (detecting overall functional anomalies).

[0246] Where: is the output feature after feature fusion of the layer encoding block in Step 2.12.

[0247] Step 3.2: Perform global average pooling (GAP) operation on to extract channel statistics:

[0248]

[0249] Where: i, j are spatial dimension coordinates (corresponding to positions on the feature map, ranging from 1×1 to 32×32): indicates full retention of the channel dimension, that is, traversing all 768 channels at each spatial position (i, j). It is used to compress the three-dimensional feature tensor (32×32×768) along the spatial dimension into a channel description vector (768 dimensions).

[0250] Step 3.3: Implement cross-channel information integration intermediate feature for through a lightweight mapping layer:

[0251]

[0252] Among them, is the parameter of the dimension-raising matrix, which expands the channel interaction capacity, is the bias vector, and GeLU is a non-linear activation function, introducing smooth non-linearity.

[0253] Step 3.4: Generate hierarchical attention weights for each layer of intermediate feature :

[0254]

[0255] Where: is the hierarchical attention weight matrix, is the hierarchical attention bias.

[0256] Step 3.5: Normalize the hierarchical attention weights to obtain the normalized hierarchical weights. .

[0257]

[0258] Where: and , realizing dynamic feature optimization.

[0259] Step 3.6: Weightedly fuse the hierarchical features with the normalized attention weights to obtain the fused features to balance details and semantics.

[0260]

[0261] Step 3.7: Align the feature dimensions of the fused features through linear projection to obtain the final aligned features .

[0262]

[0263] Where is the learnable alignment matrix, which maps high-dimensional features to the low-dimensional semantic space. The final aligned features are used for downstream feature interaction, classification, and localization tasks.

[0264] Step 4.1: Perform feature dimensionality reduction and compression on the visual features through the convolution operation to obtain the compressed features .

[0265]

[0266] Step 4.2: Generate the spatial attention aggregation weights from the compressed features through a 3×3 convolution and normalization layer.

[0267]

[0268] Step 4.3: Aggregate the context of the compressed features with the spatial attention aggregation weights to obtain the global description vector :

[0269]

[0270] Where, i , j ∈ [1,32 ] is the spatial position, and each and Multiply them together so that the spatial attention weights are applied to each position of the feature map.

[0271] Step 4.4: Convert the global description vector The dimension is increased to 1024 through the fully connected layer, and the hidden layer features with nonlinear expression ability are enhanced through the GeLU activation function. , the expression is as follows:

[0272]

[0273] in: is a dimension-raising matrix, is the bias matrix.

[0274] Step 4.5: Use the split projection strategy to transform the 1024-dimensional hidden layer features Decoupled into 6 groups of independent subspaces, each group generates a 512-dimensional latent vector through an independent linear layer

[0275]

[0276] in, is the head projection weight of the i-th group, is the projection bias of the ith group, is the linear transformation parameter.

[0277] Step 4.6: Convert 6 groups of 512-dimensional vectors Concatenate into dynamic hint prefix matrix, and obtain dynamic hint with stable feature distribution through layer normalization. :

[0278]

[0279] Step 4.7: Call the open source large language model Acquire industrial semantically enhanced knowledge:

[0280]

[0281] in: For large language models The temperature parameter controls the randomness of the generation. The larger the value, the more random the generation. Select the five most confident defect descriptions from the model output. It is a sampling strategy that limits the probability distribution of the generated results, that is, only words with a cumulative probability of 0.9 are considered. The prompt is "As an industrial quality inspection expert, list [k] types of physical defect types that may occur in the [CLASS] category during the manufacturing process, including defect morphology, location, and typical dimensions".

[0282] Among them: Represents the input category identifier (such as 15 industrial objects like "transistor", "leather" in the MVTec AD dataset), which is a semantic variable of the prompt template, guiding the large model to generate defect knowledge strongly related to the target category.

[0283] Step 4.8: Concatenate the dynamic prefix , the class name, and the "flawless" semantics to obtain the normal prompt template .

[0284]

[0285] Step 4.9: Concatenate the dynamic prefix , the class name, and the LLM semantics to obtain the abnormal prompt template

[0286]

[0287] Among them Is the statement concatenation operation, concatenating the cross-modal dynamic prompt prefix, category name, and semantic enhancement knowledge.

[0288] Step 4.10: Send the normal prompt template And the abnormal prompt template Into the CLIP text encoder to obtain the normal semantic projection And the abnormal semantic projection .

[0289]

[0290] Among them: Is 512, and K is the number of semantic enhancement items 5 in Step 4.7.

[0291] Step 5.1 Let

[0292] The aligned visual features , positive and negative text features Are respectively cross-modally semantically aligned through the normalization layer and the linear layer to obtain the triple for attention calculation:

[0293]

[0294] Among them: Is the learnable visual query attention weight, is the attention weight for normal text keys, is the attention weight for abnormal text keys, is the attention weight for normal text values, is the attention weight for abnormal text values, is the normalization operation to stabilize the feature distribution.

[0295] Step 5.2 Calculate the adaptive attention weights for each spatial position of the visual query vector : ( h , w ) ∈ [1,32 ] where:

[0296]

[0297] is the adaptive positive attention weight, is the adaptive negative attention weight, is the spatial position, is the scaling factor.

[0298] Step 5.3 Feed into to calculate the attention interaction between the visual-driven text and the normal and abnormal texts , and .

[0299]

[0300]

[0301] Step 5.4: Obtain the non-linear high-order semantic features from the attention interaction-optimized abnormal text features through the non-linear mapping layer .

[0302]

[0303]

[0304] where is the weight of the high-order semantic expansion layer, is the weight of the high-order semantic compression layer.

[0305] Step 5.5 Project the defect description generated in Step 4.7 into the semantic space through the CLIP text encoder to generate the domain knowledge bias term :

[0306]

[0307] Step 5.6 Add the non-linear high-order semantic features to the domain knowledge bias to avoid introducing noise due to modal interaction differences. Finally, normalize to obtain semantically complete abnormal text features :

[0308]

[0309] Step 5.7 At the same time, perform convolutional transformation on the aligned visual features to generate visual query vectors aligned with the text modality :

[0310]

[0311] Step 5.8 Copy and expand the optimized text features along the spatial dimension to generate key vectors aligned with the visual feature space .

[0312]

[0313] Step 5.9 Feed , into the convolutional activation gating mechanism to enhance local abnormal responses:

[0314]

[0315] Among them, identifies the channel features of the concatenation and to generate spatial gating weights through convolution and non-linear activation functions .

[0316] Step 5.10 Project the optimized text features separately into 4 heads to obtain :

[0317]

[0318] At the same time, project , to obtain:

[0319] Among them: is the projection matrix for each head.

[0320] Step 5.11 Feed , , Feed into four - head attention for fusion to obtain a multi - head fusion vector :

[0321]

[0322]

[0323] Step 5.12 Residually weight and fuse the spatial gating weights with to obtain the optimized visual features :

[0324]

[0325] where is the gated Hadamard product, realizing regional attention focusing. Realize spatial weighting of the attention result, suppress non - significant regions, and strengthen the response of abnormal regions. The gated weighted additional residual connection (connecting the original visual features ) further strengthens the response of abnormal regions.

[0326] Step 6.1: Feed the visual features , normal text features , and abnormal text features into the cross - modal contrast loss: By comparing the feature similarities of image - normal text pairs and image - abnormal text pairs, drive cross - modal feature alignment.

[0327]

[0328] where is the cosine similarity, is the temperature coefficient, controlling the distribution smoothness, is the number of normal samples 4, is the batch size 16. By comparing the similarities between image features and normal / abnormal text features, force the alignment of normal image features with normal text features, while being far from abnormal text features, enhancing cross - modal feature distinctiveness.

[0329] Step 6.2: Feed the optimized visual features , text features into the feature distribution alignment loss function:

[0330]

[0331] where: is the set of randomly sampled spatial positions ( ). It is a learnable MLP projection head. The distribution consistency between the interactive visual features and the interactive text features is constrained to maintain the stability of cross-modal interaction.

[0332] Step 6.3: Feed the normal text features and the abnormal text features into the explicit text boundary constraint loss:

[0333]

[0334] where: is the minimum margin threshold determined by experimental tuning , which forces a minimum margin between the normal and abnormal text features to avoid semantic confusion.

[0335] Step 6.4: The total loss is :

[0336]

[0337] where: = 1.0 is the boundary loss weight to strengthen the separation of normal-abnormal features, = 0.5 is the alignment loss weight to ensure the semantic consistency between vision and text.

[0338] Step 6.5: Define the optimizer and training loop: Initialize the StableAdamW optimizer, set the learning rate lr to 1×e −3 , and the weight decay wd to 1×e −4 , for 200 epochs.

[0339] Step 6.6: Forward propagation: Sequentially execute Steps 2-5 to generate the visual feature image features , the normal text features , the abnormal text features , and the optimized visual features and the optimized text features .

[0340] Step 6.7: Backward propagation and model saving: Substitute the forward propagation features in Step 6.6 into the total loss , and backpropagate to update the parameters of the industrial anomaly detection model. Complete the iteration of 200 training batches in sequence. Save the parameters of the trained dual-path visual encoder, multi-level feature adaptive fusion adapter, cross-modal dynamic prompt embedder, and cross-modal interaction module for loading in the test phase.

[0341] Step 7.0: Load the trained industrial anomaly detection model, freeze all trainable parameters, and switch to the inference mode. Load the test set data and execute Steps 2 - 5 to obtain the optimized multi-level visual features and the optimized text features .

[0342] Step 7.1: Unfold the cross-modal optimized visual features along the spatial dimension to obtain spatial features :

[0343]

[0344] Step 7.2: Calculate the cosine similarity between the visual features and the text anomaly features and perform spatial reshaping to obtain the initial anomaly response map :

[0345]

[0346]

[0347] Step 7.3: Upsample the anomaly response map to 512×512 through bilinear upsampling and obtain the dynamic learnable weight parameters through average pooling and layer normalization . .

[0348]

[0349] Step 7.4: Perform weighted fusion on the initial anomaly response map and the weight parameters to obtain the final anomaly response map .

[0350]

[0351] Step 7.5: Calculate the image-level anomaly score:

[0352] , where represents the global score and represents the local score.

[0353] Among them: combines the maximum pooling value of the global similarity and the local similarity, and is the cosine similarity.

[0354] On the MVTec AD validation set, if the image-level anomaly score > 0.85, it is determined as an anomaly.

[0355] Step 7.6: Feed into Perform pixel-level positioning:

[0356]

[0357] Where: Threshold ( is the mean, is the standard deviation) to generate an anomaly mask and achieve pixel-level positioning.

[0358] Example 3:

[0359] This example introduces a computer-readable storage medium with a computer program stored thereon. When the computer program is executed by a processor, it implements a few-shot industrial anomaly detection method based on cross-modal adaptive interaction as described in any one of Example 1.

[0360] Example 4:

[0361] This example introduces a computer device, including:

[0362] A memory for storing instructions.

[0363] A processor for executing the instructions, enabling the computer device to perform the operations of a few-shot industrial anomaly detection method based on cross-modal adaptive interaction as described in any one of Example 1.

[0364] Example 5:

[0365] This example describes conducting experiments on the authoritative benchmark dataset MVTec AD. The core research goal of this dataset is to train an industrial anomaly detection model based on normal samples to achieve accurate identification and classification of abnormal states in test samples. The experimental method adopts a per-class training and testing strategy, that is, randomly selecting 4 normal sample images from the training set for few-shot model training for each class, and immediately testing after training to complete the detection tasks for all 15 classes in turn. Using the StableAdamW optimizer, the learning rate is 1×e −3 , and the weight decay is 1×e −4 , running for 200 epochs and on an NVIDIA GeForce RTX 2080ti GPU. The evaluation metrics include image-level AUROC (Area Under the Receiver Operating Characteristic Curve) to measure the classification ability of the model to distinguish normal / abnormal images; pixel-level AUROC to evaluate the accuracy of anomaly localization. As shown in Table 1:

[0366] Table 1 is the comparison table of MVTec AD ablation experiments under 4 samples

[0367]

[0368] Through the above process, this embodiment achieves an image-level AUROC of 96.4% and a pixel-level AUROC of 96.9% on the MVTec AD dataset, verifying the efficiency and practicality of this method in industrial anomaly detection.

[0369] The method of the present invention designs dual-path block coding and multi-scale fusion for high-resolution images (512×512), balancing computational efficiency and detail retention; only through few-shot training (4 normal samples) and multi-modal dynamic prompt generation (in the past, most were only visual single-modal or multi-modal that only fused static text prompts), reducing the dependence on industrial manually labeled data; combining LLM domain knowledge injection and cross-modal bidirectional interaction mechanism, adapting to the semantic requirements of diverse working conditions, and enhancing the detection robustness of complex defects.

[0370] Among them, each module of the industrial anomaly detection model, such as Figure 2 shown, has the following advantages:

[0371] Dual-path visual encoder: Improves the CLIP model, fuses global QKV attention and local spatial-channel mixed attention (SCMA), takes into account both global semantic understanding of images and local defect area focusing, and solves the problem of detecting subtle anomalies in complex backgrounds.

[0372] Multi-level feature adaptive fusion adapter: Dynamically weights and fuses visual features at different levels in the encoder (shallow details, middle-level transitions, deep semantics), balances multi-scale information through hierarchical attention mechanism, and enhances the robustness of feature expression.

[0373] Cross-modal dynamic prompt embedder: Generates dynamic text prompt prefixes based on visual feature context, and calls a large language model (LLM) to inject domain knowledge, constructs refined normal / abnormal text templates, and enhances semantic adaptation ability.

[0374] Cross-modal interaction module: Through visual→text semantic modulation and text→visual gating mechanism, realizes bidirectional feature optimization interaction, and strengthens the response to abnormal areas and semantic consistency.

[0375] Multi-objective joint optimization loss function: Combines cross-modal contrast loss, feature distribution alignment loss and explicit text boundary constraint loss, jointly optimizes model parameters, and ensures the discriminability of features under few-shot conditions.

[0376] Multi-scale anomaly decision and scoring: Fuses multi-level feature similarity maps, combines global and local scoring, and realizes image-level anomaly determination and pixel-level defect localization.

[0377] By improving the structure of the visual encoder, the present invention introduces a global and local dual-path encoder, which retains the global semantic association of the image while enhancing the focusing ability on subtle abnormal regions; designs a multi-level feature adaptive fusion adapter to dynamically balance the details and semantic information of visual features at different levels, and constructs a cross-modal instance adaptive prompt embedding technology to align the visual features with the dynamically generated text prompts, improving the accuracy of modal interaction. On this basis, combined with the domain knowledge injection strategy of the large language model (LLM), a refined and scene-adaptive dynamic text prompt template is generated to replace the traditional static template, effectively adapting to the semantic requirements of diverse working conditions. Through the joint optimization strategy of cross-modal contrast learning, feature distribution alignment and explicit text boundary constraint, the present invention can significantly enhance the model's ability to distinguish normal and abnormal features under the condition of only requiring a small number of normal samples, reduce the dependence on labeled data, provide a high-precision, low-cost and quickly deployable industrial anomaly detection solution for intelligent manufacturing, and promote the industrial application of few-shot learning and cross-modal technology.

[0378] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A few-shot industrial anomaly detection method based on cross-modal adaptive interaction, characterized in that: Specifically, it includes: Obtain a few-shot dataset; Input the few-shot dataset into the dual-path vision encoder to obtain layers of dual-path residual fusion features; Input the dual-path residual fusion features of multiple layers in the layer dual-path residual fusion features into the multi-level feature adaptive fusion adapter to obtain the aligned final features; Input the aligned final features into a cross-modal dynamic prompt embedder to obtain normal semantic projections and abnormal semantic projections; Input the aligned final features, normal semantic projections, and abnormal semantic projections into a cross-modal interaction module to obtain optimized text features and optimized visual features; Calculate the total loss based on the aligned final features, normal semantic projections, abnormal semantic projections, optimized visual features, and optimized text features, and use the total loss to update the parameters of the industrial anomaly detection model to obtain a trained industrial anomaly detection model; Input the industrial image to be detected into the trained industrial anomaly detection model to obtain the aligned final features, abnormal semantic projections, optimized visual features, and optimized text features for judging the anomaly situation of the industrial image to be detected; The step of inputting the aligned final features into a cross-modal dynamic prompt embedder to obtain normal semantic projections and abnormal semantic projections specifically includes: Obtain a global description vector from the aligned final features, elevate the dimension of the global description vector through a fully connected layer, and obtain hidden layer features through an activation function; adopt a split-head projection strategy to decouple the hidden layer features to generate multiple groups of hidden vectors; splice the multiple groups of hidden vectors into a dynamic prompt prefix matrix and obtain a dynamic prompt prefix through layer normalization; call a large language model to obtain industrial semantic enhanced knowledge to get a defect description, splice the dynamic prompt prefix, category name, and the semantics of "flawless" to get a normal prompt template, splice the dynamic prompt prefix, category name, and industrial semantic enhanced knowledge to get an abnormal prompt template; input the normal prompt template and the abnormal prompt template into a text encoder to obtain normal semantic projections and abnormal semantic projections.

2. The few-shot industrial anomaly detection method based on cross-modal adaptive interaction according to claim 1, wherein: It also includes: Locate the abnormal situation based on the aligned final features, normal semantic projections, abnormal semantic projections, optimized visual features, and optimized text features of the industrial image to be detected.

3. A few-shot industrial anomaly detection method based on cross-modal adaptive interaction according to claim 1 or 2, characterized in that: The few-shot dataset is randomly selecting K normal samples from each type of dataset as the few-shot training set , , which are normal samples.

4. The few-shot industrial anomaly detection method based on cross-modal adaptive interaction according to claim 1 or 2, characterized in that: Inputting the few-shot dataset into the dual-path vision encoder to obtain layers of dual-path residual fusion features, specifically including: Step 2.1: The normal samples in the few-shot dataset are sequentially subjected to convolutional block division to obtain block features ; Step 2.2: Perform layer normalization on the chunk features sequentially in the channel dimension to obtain the initial chunk embedding features ; Step 2.3: Embed the initial block into the feature and input it into the global context attention path, and generate query vectors for the first-layer attention calculation through three groups of independent linear projections , key vectors and value vectors ; Step 2.4: Calculate the first-layer attention features , and the expression is as follows: ; Among them, is a normalization operation, is a scaling factor; Step 2.5: Take the first-layer attention features and input them into a feed-forward network to expand the feature dimension and obtain intermediate-layer features and the first-layer QKV attention , and the expression is as follows: ; ; Among them, is the weight parameter of the feedforward network expansion layer, is the bias parameter of the feedforward network expansion layer, is the weight parameter of the feedforward network compression layer, is the bias parameter of the feedforward network compression layer, is the activation function; Step 2.6: Embed the initial block into the features Input space-channel hybrid attention path, generating the first-layer spatial weight map through lightweight single-channel convolution ; Step 2.7: Calculate the initial block embedding features The first-layer global average pooling value ; Step 2.8: The first-layer global average pooling value is projected to a higher dimension through a linear layer and passed through an activation function to obtain intermediate features ; Step 2.9: The intermediate feature is restored by dimensionality reduction through a linear projection layer and the first-layer channel weights are obtained through an activation function ; Step 2.10: Take the first-layer spatial weight map , the first-layer channel weights and the initial block embedding features to output the first-layer fused features through hybrid gating weighting ; Step 2.11: Calculate the first-layer dual-path residual fusion features based on the first-layer QKV attention , the features after the first-layer fusion and the initial block embedding features , with the expression as follows: The expression is as follows: ; Step 2.12 Repeat steps 2.4 - 2.11 to calculate respectively of layer double - path residual fusion features , and the expression is as follows: 。 5. A few-shot industrial anomaly detection method based on cross-modal adaptive interaction according to claim 1 or 2, characterized in that: The step of inputting the dual-path residual fusion features of multiple layers in the layer dual-path residual fusion features into a multi-level feature adaptive fusion adapter to obtain the finally aligned features, specifically includes: Step 3.1: Extract from the double-path residual fusion features of the th , , layers as the hierarchical features , , where ; , , ; Step 3.2: Perform global average pooling on the hierarchical features respectively to obtain average pooling values , where , , ; Step 3.3: Obtain intermediate features from the average pooling value through a lightweight mapping layer , and the expression is as follows: ; Among them, is the dimension-raising matrix parameter, is the bias vector, and GeLU is the activation function; Step 3.4: Calculate the intermediate features of each layer of the hierarchical attention weights , and the expression is as follows: ; Wherein: is the hierarchical attention weight matrix, is the hierarchical attention bias; Step 3.5: Perform weight normalization on the hierarchical attention weights to obtain the hierarchical weight normalized weights ; Step 3.6: Hierarchical features are weighted and fused with the normalized attention weights to obtain the fused features ; ; Step 3.7: Align the dimensionality of the fused features through linear projection to obtain the final aligned features ; ; Wherein: is a learnable alignment matrix.

6. The few-shot industrial anomaly detection method based on cross-modal adaptive interaction according to claim 1 or 2, characterized in that: The step of inputting the aligned final features into a cross-modal dynamic prompt embedder to obtain normal semantic projections and abnormal semantic projections specifically includes: Step 4.1: Final Aligned Features Obtain compressed features through convolution operation ; Step 4.2: Generate spatial attention aggregation weights for the compressed features through convolution and normalization layers ; Step 4.3: Compress the features and the spatial attention aggregation weights to perform context aggregation to obtain a global description vector ; Step 4.4: Upscale the global description vector through a fully connected layer, and obtain the hidden layer features through an activation function ; Step 4.5: Adopt a split projection strategy to decouple the hidden layer features into S groups of independent subspaces, and each group generates a hidden vector through an independent linear layer , where , and the expression is as follows: ; Among them, is the head projection weight of the i-th component, is the head projection bias of the i-th component, is the linear transformation parameter; Step 4.6: Concatenate the S-group hidden vectors to form a dynamic prompt prefix matrix, and obtain the dynamic prompt through layer normalization ; Step 4.7: Call the large language model to obtain industrially semantically enhanced knowledge , and the expression is as follows: ; Wherein: is a large language model, are the Z defect descriptions with the highest confidence selected from the output of the large language model, represents the input category identifier; Step 4.8: Concatenate the dynamic prefix , the class name, and the "flawless" semantics to obtain a normal prompt template ; ; Step 4.9: Concatenate the dynamic prefix , class name, and industrial semantic enhancement knowledge to obtain an exception prompt template ; ; Wherein: is a statement splicing operation; Step 4.10: Send the normal prompt template and the abnormal prompt template into the text encoder to obtain the normal semantic projection and the abnormal semantic projection .

7. A few-shot industrial anomaly detection method based on cross-modal adaptive interaction according to claim 1 or 2, characterized in that: The step of inputting the aligned final features, normal semantic projections, and abnormal semantic projections into a cross-modal interaction module to obtain optimized text features and optimized visual features specifically includes: Step 5.1: Let be the positive and negative text features, and the aligned final features , positive and negative text features are respectively cross-modally semantically aligned through a normalization layer and a linear layer to obtain a visual query attention vector for attention calculation, a normal text key attention vector , an abnormal text key attention vector , a normal text value attention vector and an abnormal text value attention vector , and the expressions are as follows: ; Wherein: is the learnable visual query attention weight, is the normal text key attention weight, is the abnormal text key attention weight, is the normal text value attention weight, is the abnormal text value attention weight, is the normalization operation, is the normal semantic projection, is the abnormal semantic projection; Step 5.2: Calculate the adaptive positive attention weights for each spatial position of the visual query attention vector and the adaptive negative attention weights , and the expressions are as follows: ; Wherein: is a spatial position, is a scaling factor; Step 5.3: According to the adaptive positive attention weight , the adaptive negative attention weight , calculate the normal text features after attention interaction optimization , the abnormal text features after attention interaction optimization , and the expression is as follows: ; ; Among them, represents the total height value, and b represents the total width value; Step 5.4: The abnormal text features after optimizing the attention interaction Obtain non-linear high-order semantic features through a non-linear mapping layer , and the expression is as follows ; ; Wherein: is the weight of the high-order semantic expansion layer, is the bias of the high-order semantic expansion layer; is the weight of the high-order semantic compression layer, the bias of the high-order semantic compression layer; Step 5.5: Project the industrial semantic enhanced knowledge The generated defect description is projected into the semantic space through a text encoder to generate a domain knowledge bias term ; Step 5.6: After adding the non-linear high-order semantic features and the domain knowledge bias term and normalizing, obtain the anomaly text features with complete semantics ; Step 5.7: Convolve the aligned final features to generate a visual query vector aligned with the text modality '; ; Step 5.8: Copy and expand the optimized text features along the spatial dimension to generate a key vector aligned with the visual feature space ; ; Among them, represents the optimized text feature, and let ; Step 5.9: Generate spatial gating weights for the visual query vector aligned with the text modality and the key vector aligned with the visual feature space through convolution and a non-linear activation function ; Step 5.10: Divide into multiple heads and project to obtain a multi-headed text value vector , divide , and project to obtain a multi-headed visual query vector , a multi-headed text key vector . The expression is as follows: ; ; ; Wherein: is the query projection matrix for each head, is the key projection matrix for each head, is the value projection matrix for each head; Step 5.11: Combine , , by inputting them into the head attention to obtain the multi-head fusion vector . The expression is as follows: ; ; Among them, is the multi-head output fusion weight; Step 5.12: Residually weight and fuse the spatial gating weights with the multi-head fusion vector and the aligned final feature to obtain the optimized visual feature , and the expression is as follows: ; Among them, is the gated product.

8. A few-shot industrial anomaly detection method based on cross-modal adaptive interaction according to claim 1 or 2, characterized in that: The step of calculating the total loss based on the aligned final features, normal semantic projections, abnormal semantic projections, optimized visual features, and optimized text features, and using the total loss to update the parameters of the industrial anomaly detection model to obtain a trained industrial anomaly detection model specifically includes: Step 6.1: Initialize the optimizer, set the learning rate, weight decay, and the number of iterations, and perform forward propagation in Steps 2 - 5 to generate the final aligned features. , normal semantic projection and abnormal semantic projection , the optimized visual features and the optimized text features ; Step 6.2: According to the aligned final features , normal semantic projection and abnormal semantic projection calculate the cross-modal contrastive loss , and the expression is as follows: ; Among them, is the cosine similarity, is the temperature coefficient, is the number of normal samples, is the batch, 、 、 respectively represent the final feature, normal semantic projection, and abnormal semantic projection after alignment of the i-th normal sample, and e is the natural constant; Step 6.3: Calculate the feature distribution alignment loss function according to the optimized visual features , the optimized text features , as follows: The expression is as follows: ; Wherein: is a set of spatially sampled positions, is a visual alignment projection head is a text alignment projection head, represents the feature vector of the optimized visual feature at the spatial position (h, w), represents the norm of 2; Step 6.4: According to the normal semantic projection and the abnormal semantic projection calculate the explicit text boundary constraint loss , and the expression is as follows: ; Wherein: is the minimum interval threshold value, represents the maximum value function; Step 6.5: Calculate the total loss according to the cross-modal contrastive loss , the feature distribution alignment loss function , and the explicit text boundary constraint loss , and the expression is as follows: ​ ; Wherein: is the boundary loss weight, is the alignment loss weight; Step 6.6: Update the parameters of the industrial anomaly detection model through backpropagation, complete the number of training iterations in sequence, and save the parameters of the trained dual-path visual encoder, multi-level feature adaptive fusion adapter, cross-modal dynamic prompt embedder, and cross-modal interaction module to obtain a trained industrial anomaly detection model.

9. A few-shot industrial anomaly detection method based on cross-modal adaptive interaction according to claim 1 or 2, characterized in that: The step of inputting the industrial image to be detected into the trained industrial anomaly detection model to obtain the aligned final features, abnormal semantic projections, optimized visual features, and optimized text features for judging the anomaly situation of the industrial image to be detected specifically includes: Step 7.1: Input the industrial image to be detected into the trained industrial anomaly detection model to obtain the final aligned features , anomaly semantic projection , optimized visual features and optimized text features ; Step 7.2: Unfold the optimized visual features along the spatial dimension to obtain spatial features ; Step 7.3: Calculate spatial features and the abnormal semantic projection to obtain the cosine similarity and perform spatial reshaping to generate an initial abnormal response map , and the expression is as follows: ; ; Step 7.4: Take the initial anomaly response map Perform upsampling, and obtain dynamically learnable weight parameters through average pooling and layer normalization ; Step 7.5: Combine the initial anomaly response map with the weight parameter through weighted fusion to obtain the final anomaly response map ; Step 7.6: Calculate the image-level anomaly score based on the finally aligned features and the optimized text features The expression is as follows: , as follows: ; Wherein: represents the global similarity and the maximum pooling value parameter of the local similarity, is the cosine similarity; Step 7.7: When the image-level anomaly score is greater than the threshold, it is determined as an abnormal situation.

10. A few-shot industrial anomaly detection method based on cross-modal adaptive interaction according to claim 2, characterized in that: Based on the aligned final features of the industrial image to be detected, the normal semantic projection, abnormal semantic projection, optimized visual features, and optimized text features are used to locate abnormal situations, specifically including: Step 8.1: Unfold the optimized visual features along the spatial dimension to obtain spatial features ; Step 8.2: Calculate the spatial features and the abnormal semantic projection to obtain the cosine similarity and perform spatial reshaping to generate an initial abnormal response map , and the expression is as follows: ; ; Step 8.3: Take the initial anomaly response map Through upsampling, and obtain the dynamically learnable weight parameters through average pooling and layer normalization ; Step 8.4: Combine the initial anomaly response map with the weight parameter through weighted fusion to obtain the final anomaly response map ; Step 8.5: According to the final abnormal response diagram , calculate the abnormal mask , and obtain the abnormal situation positioning; ; Among them, is the threshold value, and 1 is the pixel coordinate of the abnormal situation.

Citation Information

Patent Citations

  • Industrial anomaly detection method and system based on fine-grained text prompt feature engineering

    CN118568650A

  • Zero sample image anomaly detection method and device

    CN119130931A