Small-sample industrial anomaly detection method based on cross-modal adaptive interaction
By adopting a cross-modal adaptive interaction method in industrial anomaly detection, the features are extracted using a dual-path vision encoder and a multi-level feature adaptive fusion adapter, and the features are optimized through a cross-modal dynamic prompt embedder and interaction module, the problems of low efficiency of abnormal feature extraction and cross-modal interaction in few-sample scenarios are solved, and high-precision and low-cost industrial anomaly detection are achieved.
Patent Information
- Application Number
- CN202510484103.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing industrial anomaly detection methods are difficult to efficiently extract abnormal features in scenarios with few samples, and the cross-modal interaction is inefficient, resulting in low detection accuracy and high missed detection rate.
Using a small sample industrial anomaly detection method based on cross-modal adaptive interaction, features are extracted through a dual-path vision encoder and a multi-level feature adaptive fusion adapter, and feature optimization is performed using a cross-modal dynamic prompt embedder and a cross-modal interaction module.
Effectively overcome complex background interference, improve detection sensitivity for small defects, reduce missed detection rate, improve detection accuracy, and reduce dependence on labeled data.
Smart Images

Figure CN119989247A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a few-sample industrial anomaly detection method based on cross-modal adaptive interaction, and belongs to the technical field of industrial detection. Background Art
[0002] Industrial anomaly detection is a core part of intelligent manufacturing, which aims to identify and locate defects or anomalies in products or processes through automated means. Traditional methods mostly rely on supervised learning, which requires a large number of anomaly samples to be labeled. However, anomaly samples in industrial scenarios are scarce and the labeling cost is high, making it difficult for the model to adapt to actual industrial scenarios.
[0003] In recent years, Few-Shot Anomaly Detection (FSAD) has become a research topic that has attracted much attention. FSAD aims to use limited normal samples to train the model and infer the existence of anomalies by learning the distribution characteristics of normal data. Compared with traditional methods, few-shot learning can achieve higher detection performance in scenarios with limited data, and therefore has important industrial application potential.
[0004] However, the implementation of few-shot anomaly detection still faces technical challenges, especially in how to efficiently extract anomaly features and improve detection accuracy. At the same time, multimodal vision-language models (VLM, such as CLIP: Contrastive Language-Image Pre-Training) have made significant progress in the intersection of computer vision and natural language processing. These models demonstrate strong performance in tasks such as image classification and object detection by fusing visual information (such as images) and text information (such as descriptive language). The core advantage of VLM lies in its cross-modal capability, that is, guiding visual feature extraction through text or reversely enhancing semantic understanding through visual information. Therefore, applying VLM with cross-modal capabilities to industrial anomaly detection, especially in few-shot scenarios, has significant practical value.
[0005] However, existing anomaly detection algorithms have some defects:
[0006] (1) Single-modality limitations: Existing traditional methods are mostly based on pure visual features (such as image segmentation and reconstruction errors), lack a deep understanding of defect semantics, and have difficulty distinguishing subtle anomalies in complex backgrounds.
[0007] (2) Inefficient cross-modal interaction: A few multimodal methods that combine vision and text have problems with rough feature fusion and imprecise modal alignment, resulting in insufficient response to local defects. How to effectively guide text information and visual features to interact and prompt each other, and achieve accurate identification of anomalies, remains a difficult problem that needs to be solved.
[0008] (3) Insufficient semantic information: Most existing multimodal methods use static text prompt templates (such as "a photo of [state]"), which lack the injection of refined domain knowledge and are difficult to dynamically adapt to the semantic requirements of different industrial scenarios, resulting in an increased missed detection rate. Summary of the invention
[0009] Objective: To overcome the problems in the prior art of complex background interference caused by the limitations of single-modal features, insufficient local defect response caused by inefficient cross-modal interaction, and high missed detection rate caused by weak semantic adaptation ability of static text templates, the present invention provides a few-sample industrial anomaly detection method based on cross-modal adaptive interaction.
[0010] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is:
[0011] In the first aspect, a few-sample industrial anomaly detection method based on cross-modal adaptive interaction specifically includes:
[0012] Get a few-shot dataset.
[0013] The few-shot dataset is fed into the dual-path visual encoder to obtain Layer dual-path residual fusion features.
[0014] Will The dual-path residual fusion features of multiple layers in the layer dual-path residual fusion feature are input into the multi-level feature adaptive fusion adapter to obtain the final aligned features.
[0015] The aligned final features are input into the cross-modal dynamic cue embedder to obtain normal semantic projection and abnormal semantic projection.
[0016] The aligned final features, normal semantic projections and abnormal semantic projections are input into the cross-modal interaction module to obtain optimized text features and optimized visual features.
[0017] The total loss is calculated based on the aligned final features, normal semantic projection, abnormal semantic projection, optimized visual features and optimized text features, and the total loss is used to update the parameters of the industrial anomaly detection model to obtain a trained industrial anomaly detection model.
[0018] The industrial image to be detected is input into the trained industrial anomaly detection model to obtain the aligned final features, anomaly semantic projection, optimized visual features and optimized text features, which are used to judge the anomaly of the industrial image to be detected.
[0019] The abnormal situation is located according to the aligned final features of the industrial image to be detected, the normal semantic projection and the abnormal semantic projection, the optimized visual features and the optimized text features.
[0020] In a second aspect, a computer-readable storage medium stores a computer program, which, when executed by a processor, implements a few-sample industrial anomaly detection method based on cross-modal adaptive interaction as described in the first aspect.
[0021] According to a third aspect, a computer device includes:
[0022] Memory, used to store instructions.
[0023] The processor is used to execute the instructions so that the computer device performs the operations of the few-sample industrial anomaly detection method based on cross-modal adaptive interaction as described in the first aspect.
[0024] Beneficial effects: The invention provides a method for industrial anomaly detection based on cross-modal adaptive interaction with a small number of samples. By introducing a cross-modal bidirectional adaptive interaction mechanism and a small number of sample optimization strategy, the invention shows significant beneficial effects in the field of industrial anomaly detection. Specifically, the method includes the following:
[0025] 1. Based on the dual-path attention visual encoder and the multi-level feature adaptive fusion adapter, the model can simultaneously capture the global semantic associations and local subtle abnormal features of the image, effectively overcome complex background interference, and improve the detection sensitivity of tiny defects.
[0026] 2. By embedding dynamic text prompts and injecting domain knowledge of large language models, the problem of insufficient semantic adaptation of traditional static templates is solved, and refined, scenario-adaptive exception description templates are generated, which significantly reduces the missed detection rate and enhances the accuracy of cross-modal interaction.
[0027] 3. The bidirectional cross-modal interactive optimizer combines the multi-objective joint optimization loss function to achieve efficient alignment of visual and text features and enhancement of abnormal features. Under the condition of only 4 normal reference samples, the model can still maintain high robustness, reduce dependence on labeled data, and significantly reduce industrial quality inspection costs.
[0028] 4. The multi-scale anomaly scoring and positioning module can not only provide image-level anomaly detection and judgment, but also locate pixel-level anomalies by fusing global and local feature responses, supporting high-precision defect detection and real-time decision-making, providing a low-cost, high-precision and fast-deployment solution for smart manufacturing scenarios, and promoting the actual industrial application of few-sample learning and cross-modal technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1The present invention is a flowchart of a method for detecting industrial anomalies with a small number of samples based on cross-modal adaptive interaction.
[0030] Figure 2 It is a structural schematic diagram of the industrial anomaly detection model of the present invention. DETAILED DESCRIPTION
[0031] The following is a clear and complete description of the technical solutions in the examples of the present invention in conjunction with the accompanying drawings in the examples of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the protection scope of the present invention.
[0032] The present invention will be further described below in conjunction with specific embodiments.
[0033] Embodiment 1:
[0034] This embodiment introduces a few-sample industrial anomaly detection method based on cross-modal adaptive interaction. Figure 1 As shown, specifically including:
[0035] Step 1: Get a few-shot dataset.
[0036] Furthermore, the few-sample data set specifically includes:
[0037] K normal samples of each category are selected from the authoritative benchmark dataset for industrial anomaly detection.
[0038] Among them, the authoritative benchmark dataset for industrial anomaly detection uses MVTec AD and covers high-quality images with a resolution of 1024×1024 for 15 categories of industrial products.
[0039] Furthermore, in one embodiment, K normal samples (uniformly resized to 512×512, RGB three channels) are randomly selected from each type of data set to form a few-sample training set. , , The reference images are randomly selected. This type of data is characterized by high resolution, low sample size, and missing annotations, which meets the actual needs of industrial quality inspection scenarios where abnormal samples are scarce. The output is a set of small sample reference images with uniform scale. , as input for subsequent feature extraction.
[0040] Step 2: Input the few-shot dataset into the dual-path visual encoder to obtain Layer dual-path residual fusion features.
[0041] Furthermore, the step 2 specifically includes:
[0042] Step 2.1: Convolution blocks are performed sequentially to obtain block features , the expression is as follows:
[0043]
[0044] in: is a 16×16 convolution kernel with a stride of 16 to split the input image into blocks (each block has a resolution of 16×16), i is the sample extracted for each class, ranging from 1-K.
[0045] Step 2.2: Sequentially block features in the channel dimension Execution layer normalization , get the initial block embedding feature , the expression is as follows:
[0046]
[0047] Step 2.3: Embedding the initial block into features Input to the global context attention path, generate the triple vector of the first layer attention calculation through three sets of independent linear projections , and , the expression is as follows:
[0048]
[0049] in: is the trainable query vector parameter, is the trainable key vector parameter, is a trainable value vector parameter.
[0050] Step 2.4: Calculate the first-level attention features , the expression is as follows:
[0051]
[0052] in, For normalization operation, It is a scaling factor to prevent the inner product from being too large and causing gradient instability. T represents the transposed matrix.
[0053] Step 2.5: The first-level attention features Input feedforward network to expand feature dimension to obtain intermediate layer features 、 First layer QKV attention , the expression is as follows:
[0054]
[0055]
[0056] in, is the weight parameter of the feedforward network extension layer, is the bias parameter of the feedforward network extension layer, is the weight parameter of the compression layer of the feedforward network, is the bias parameter of the compression layer of the feedforward network.
[0057] Step 2.6: Embedding the initial block into features Input space-channel mixed attention path ( ), the first layer of spatial weight map is generated by lightweight single-channel convolution , the expression is as follows:
[0058]
[0059] in: for The convolution kernel has an output channel of 1. is the activation function.
[0060] Step 2.7: Calculate initial block embedding features The first layer global average pooling value , the expression is as follows:
[0061]
[0062] Step 2.8: Pool the first layer globally averaged Through linear layer dimensional projection and activation function, nonlinear intermediate quantities are introduced to fit complex channel relationships to obtain intermediate features.
[0063]
[0064] in: is the linear transformation dimension-raising parameter, is the linear transformation bias parameter.
[0065] Step 2.9: Transform the intermediate features The first layer channel weights are obtained through dimensionality reduction and restoration through the linear projection layer and activation function , the expression is as follows:
[0066]
[0067] in: is the dimension reduction matrix parameter of the projection layer, It is the bias parameter of the projection layer dimensionality reduction.
[0068] Step 2.10: Transform the first layer of spatial weight map , the first layer channel weight Embedding features with the initial block Output the first layer of fused features through mixed gating weights , the expression is as follows:
[0069]
[0070] in, is the gated Hadamard product, which realizes regional gated attention focus. It is an element-by-element multiplication to achieve channel dimension scaling.
[0071] Step 2.11: Based on the first layer of QKV attention , the features after the first layer fusion and the initial block embedding features Calculate the first layer of dual-path residual fusion features , the expression is as follows:
[0072]
[0073] Step 2.12 Repeat steps 2.4 to 2.11 to calculate of Layer dual-path residual fusion feature , the expression is as follows:
[0074]
[0075] Step 3: The dual-path residual fusion features of multiple layers in the layer dual-path residual fusion feature are input into the multi-level feature adaptive fusion adapter to obtain the final aligned features.
[0076] Furthermore, the step 3 specifically includes:
[0077] Step 3.1: From Layer dual-path residual fusion feature Extract the , , Dual-path residual fusion features of the layer , As a hierarchical feature ,in, , , .
[0078] Step 3.2: Layer features Perform global average pooling respectively to obtain the average pooling value ,in, , , .
[0079]
[0080] in: The channel representing each spatial position (i, j) is traversed.
[0081] Step 3.3: Average pooling value Obtain intermediate features through lightweight mapping layers , the expression is as follows:
[0082]
[0083] in, is the dimension-upgrading matrix parameter, expanding the channel interaction capacity, is the bias vector, GeLU is a nonlinear activation function, and smooth nonlinearity is introduced.
[0084] Step 3.4: Calculate the intermediate features of each layer The level attention weights , the expression is as follows:
[0085]
[0086] in: is the level attention weight matrix, It is hierarchical attention bias.
[0087] Step 3.5: Set the level attention weights Normalize the weights to get the hierarchical weight normalization weights .
[0088]
[0089] in: Represents an exponential function with a natural constant as base.
[0090] Step 3.6: Layer features And the normalized attention weight Weighted fusion to obtain fusion features .
[0091]
[0092] Step 3.7: Fusion features Align the feature dimensions through linear projection to obtain the final aligned features .
[0093]
[0094] in: is the learnable alignment matrix.
[0095] Step 4: Input the aligned final features into the cross-modal dynamic cue embedder to obtain normal semantic projection and abnormal semantic projection.
[0096] Furthermore, the step 4 specifically includes:
[0097] Step 4.1: Final features after alignment Obtaining compressed features through convolution operations , the expression is as follows:
[0098]
[0099] Step 4.2: Compress the features Generate spatial attention aggregation weights through convolution and normalization layers , the expression is as follows:
[0100]
[0101] Step 4.3: Compress the features and spatial attention weights Perform context aggregation to obtain a global description vector , the expression is as follows:
[0102]
[0103] in, is the spatial position, , Represents the compression features of each spatial position and spatial attention aggregation weights .
[0104] Step 4.4: Convert the global description vector The dimension is increased through the fully connected layer, and the hidden layer features are obtained through the activation function , the expression is as follows:
[0105]
[0106] in: is a dimension-raising matrix, is the bias matrix.
[0107] Step 4.5: Use the split projection strategy to transform the hidden layer features Decoupled into S groups of independent subspaces, each group generates a hidden vector through an independent linear layer ,in, , the expression is as follows:
[0108]
[0109] in, is the head projection weight of the i-th group, is the projection bias of the ith group, is the linear transformation parameter.
[0110] Step 4.6: Set S as the latent vector Concatenate into dynamic hint prefix matrix and get dynamic hint by layer normalization :
[0111]
[0112] Step 4.7: Call the large language model to obtain industrial semantic enhancement knowledge , the expression is as follows:
[0113]
[0114] in: For large language models, To select the Z most confident defect descriptions from the output of the large language model, Represents the category identifier of the input.
[0115] Step 4.8: Add the dynamic prefix , the class name is concatenated with the "flawless" semantics to get a normal prompt template .
[0116]
[0117] Step 4.9: Add the dynamic prefix , class names and industrial semantics enhanced knowledge Perform splicing to get the abnormal prompt template .
[0118]
[0119] in For sentence concatenation operations, cross-modal dynamic prompt prefixes, category names, and semantically enhanced knowledge are concatenated.
[0120] Step 4.10: Normal prompt template With exception prompt template Feed into the text encoder to get normal semantic projection Projection with exception semantics .
[0121]
[0122] Step 5: Input the aligned final features, normal semantic projections, and abnormal semantic projections into the cross-modal interaction module to obtain optimized text features and optimized visual features.
[0123] Furthermore, the step 5 specifically includes:
[0124] Step 5.1: Order As positive and negative text features, the final features after alignment , positive and negative text features The attention calculation triples are obtained by cross-modal semantic alignment through the normalization layer and the linear layer. , and , including the visual query attention vector of the attention calculation , normal text key attention vector , abnormal text key attention vector , normal text value attention vector and the abnormal text value attention vector , the expression is as follows:
[0125]
[0126] in: for learnable visual query attention weights, is the normal text key attention weight, is the attention weight of the abnormal text key, is the normal text value attention weight, is the attention weight of the abnormal text value, It is a normalization operation to stabilize the feature distribution.
[0127] Step 5.2: Attention vector for visual query Calculate the adaptive positive attention weight for each spatial position of , Adaptive negative attention weight , the expression is as follows:
[0128]
[0129] in: is the adaptive positive attention weight, is the adaptive negative attention weight, is the spatial position, is the scaling factor.
[0130] Step 5.3: Adaptive positive attention weights , Adaptive negative attention weight , calculate the normal text features after attention interaction optimization , abnormal text features after attention interaction optimization , the expression is as follows:
[0131]
[0132]
[0133] in, They are respectively normal text features and abnormal text features after attention interaction optimization, so that the text features are optimized by dynamically focusing on potential abnormal areas in the image.
[0134] Step 5.4: Abnormal text features after attention interaction optimization Nonlinear high-order semantic features are obtained through nonlinear mapping layers .
[0135]
[0136]
[0137] in: is the weight of the high-order semantic extension layer, Bias for the higher-level semantic extension layer. is the weight of the high-order semantic compression layer, High-level semantic compression layer bias.
[0138] Step 5.5: Enhance knowledge with industry semantics The generated defect description is projected into the semantic space through the text encoder to generate domain knowledge bias items , the expression is as follows:
[0139]
[0140] Step 5.6: Nonlinear high-order semantic features Bias with domain knowledge After addition, normalization is performed to obtain semantically complete abnormal text features. , the expression is as follows:
[0141]
[0142] Step 5.7: Align the final features Perform convolution transformation to generate a visual query vector aligned with the text modality ', the expression is as follows:
[0143]
[0144] Step 5.8: Copy and expand the optimized text features along the spatial dimension to generate a key vector aligned with the visual feature space .
[0145]
[0146] in, represents the optimized text features, .
[0147] Step 5.9: Align the visual query vector with the text modality , a key vector aligned with the visual feature space Generate spatial gating weights through convolution and nonlinear activation functions , the expression is as follows:
[0148]
[0149] Step 5.10: Divide the H-head projections to obtain multi-head text value vectors ,Will , Projecting multi-head visual query vector , multi-headed text key vector , the expression is as follows:
[0150]
[0151]
[0152]
[0153] in: is the query projection matrix for each head, is the key projection matrix for each head, The projection matrix for each head value.
[0154] Step 5.11: , , Input H head attention to fuse and get multi-head fusion vector , the expression is as follows:
[0155]
[0156]
[0157] in, is the multi-head output fusion weight.
[0158] Step 5.12: Set the spatial gating weights Fusion vector with multiple heads And the final features after alignment Perform residual weighted fusion to obtain optimized visual features :
[0159]
[0160] in, is the gated product.
[0161] Step 6: Calculate the total loss based on the aligned final features, normal semantic projection, abnormal semantic projection, optimized visual features and optimized text features, and use the total loss to update the parameters of the industrial anomaly detection model to obtain a trained industrial anomaly detection model, where the industrial anomaly detection model includes: a dual-path visual encoder connected in sequence, a multi-level feature adaptive fusion adapter, a cross-modal dynamic prompt embedder and a cross-modal interaction module.
[0162] Furthermore, the step 6 specifically includes:
[0163] Step 6.1: Initialize the optimizer, set the learning rate, weight decay, and the number of iterations, execute steps 2-5 for forward propagation, and generate the final features after alignment , normal semantic projection Projection with exception semantics , optimized visual features And the optimized text features .
[0164] Step 6.2: Based on the final features after alignment , normal semantic projection Projection with exception semantics Compute cross-modal contrast loss , the expression is as follows:
[0165]
[0166] in, is the cosine similarity, is the temperature coefficient, is the normal sample size, For batches, , , They represent the final features, normal semantic projection and abnormal semantic projection of the i-th normal sample respectively, and e is a natural constant.
[0167] Step 6.3: Based on the optimized visual features , optimized text features Calculate the feature distribution alignment loss function , the expression is as follows:
[0168]
[0169] in: is a set of randomly sampled spatial locations. Align the projection head for vision Align the drop shadow header for the text, Represents the feature vector of the optimized visual feature at the spatial position (h, w). Represents the norm of 2.
[0170] Step 6.4: Projection based on normal semantics Projection with exception semantics Compute explicit text boundary constraint loss , the expression is as follows:
[0171]
[0172] in: is the minimum interval threshold, Represents the maximum value function.
[0173] Step 6.5: Based on cross-modal contrast loss , feature distribution alignment loss function and explicit text boundary constraint loss Calculate total loss , the expression is as follows:
[0174]
[0175] in: is the boundary loss weight, is the alignment loss weight.
[0176] Step 6.6: Back-propagation updates the parameters of the industrial anomaly detection model, completes the number of training iterations in turn, saves the parameters of the trained dual-path visual encoder, multi-level feature adaptive fusion adapter, cross-modal dynamic prompt embedder and cross-modal interaction module, and obtains the trained industrial anomaly detection model.
[0177] Step 7: Input the industrial image to be detected into the trained industrial anomaly detection model to obtain the aligned final features, anomaly semantic projection, optimized visual features, and optimized text features, which are used to determine the anomaly of the industrial image to be detected.
[0178] Furthermore, the step 7 specifically includes:
[0179] Step 7.1: Input the industrial image to be detected into the trained industrial anomaly detection model to obtain the final features after alignment , Abnormal semantic projection , optimized visual features And the optimized text features .
[0180] Step 7.2: Optimize the visual features Expand along the spatial dimension to obtain spatial features , the expression is as follows:
[0181]
[0182] Step 7.3: Calculate spatial features Projection with exception semantics The cosine similarity of and spatial reshaping is performed to generate the initial abnormal response map :
[0183]
[0184]
[0185] Step 7.4: Map the initial abnormal response to Through upsampling, average pooling and layer normalization, dynamic learnable weight parameters are obtained , the expression is as follows:
[0186]
[0187] Step 7.5: Map the initial abnormal response With weight parameters Perform weighted fusion to obtain the final abnormal response map , the expression is as follows:
[0188]
[0189] Step 7.6: Based on the final features after alignment , optimized text features Calculating image-level anomaly scores , the expression is as follows:
[0190]
[0191] in: Represents the maximum pooling value parameters of global similarity and local similarity, is the cosine similarity.
[0192] Step 7.7: When image-level anomaly scoring When it is greater than the threshold, it is determined to be an abnormal situation.
[0193] Step 8: According to the aligned final features of the industrial image to be detected, the normal semantic projection and the abnormal semantic projection, the optimized visual features and the optimized text features, the abnormal situation is located.
[0194] Furthermore, the step 8 specifically includes:
[0195] Step 8.1: Optimize the visual features Expand along the spatial dimension to obtain spatial features , the expression is as follows:
[0196]
[0197] Step 8.2: Calculate spatial features Projection with exception semantics The cosine similarity of and spatial reshaping is performed to generate the initial abnormal response map :
[0198]
[0199]
[0200] Step 8.3: Map the initial abnormal response Through upsampling, average pooling and layer normalization, dynamic learnable weight parameters are obtained , the expression is as follows:
[0201]
[0202] Step 8.4: Map the initial abnormal response to With weight parameters Perform weighted fusion to obtain the final abnormal response map , the expression is as follows:
[0203]
[0204] Step 8.5: Based on the final abnormal response graph , calculate the anomaly mask , get the abnormal situation location.
[0205]
[0206] in, is the threshold value, and 1 is the pixel coordinate of the abnormal situation.
[0207] Embodiment 2:
[0208] This embodiment introduces an embodiment of a few-sample industrial anomaly detection method based on cross-modal adaptive interaction, which specifically includes:
[0209] Step 1.1: Obtain the MVTec AD public dataset for industrial anomaly detection, which contains high-quality color images of 15 industrial detection object categories, totaling 5354 samples.
[0210] Step 1.2: Randomly select from each category in turn =4 samples constitute a few-sample normal image sample set, which is used as the training set. , is a randomly selected reference image. The training set only contains normal samples without defects.
[0211] Step 2.1: Output from step 1 Convolution blocks are performed sequentially to obtain block features , the expression is as follows:
[0212]
[0213] in: is a 16×16 convolution kernel with a stride of 16 to split the input image into blocks (each with a resolution of 16×16), i is the sample extracted for each class, ranging from 1 to 4.
[0214] Step 2.2: Sequentially block features in the channel dimension Execution layer normalization , get the initial block embedding feature , the expression is as follows:
[0215]
[0216] Step 2.3: Embedding the initial block into features Input to the global context attention path, three sets of independent linear projections are used to generate the triple vector for the first layer of attention calculation:
[0217]
[0218] in: is the trainable query vector parameter, is the trainable key vector parameter, is a trainable value vector parameter.
[0219] Step 2.4: Calculate the first-level attention features , the expression is as follows:
[0220]
[0221] in, For normalization operation, It is a scaling factor to prevent the inner product from being too large and causing gradient instability. T represents the transposed matrix.
[0222] Step 2.5: The first-level attention features Input feedforward network to expand feature dimension to obtain intermediate layer features 、 First layer QKV attention :
[0223]
[0224]
[0225] in, is the weight parameter of the feedforward network extension layer, is the bias parameter of the feedforward network extension layer, is the weight parameter of the compression layer of the feedforward network, Compress the bias parameters of the feed-forward network layer to enhance the nonlinear expression ability of the features.
[0226] is a nonlinear activation function:
[0227]
[0228] Step 2.6: Embed the initial block into the feature at the same time Feed into the spatial-channel hybrid attention path ( ), the first layer of spatial weight map is generated by lightweight single-channel convolution , the expression is as follows:
[0229]
[0230] in: for The convolution kernel has an output channel of 1. is the activation function.
[0231] Step 2.7: Calculate initial block embedding features The first layer global average pooling value , each The feature map is compressed into a scalar to compress the spatial dimension and generate a channel description vector.
[0232]
[0233] Step 2.8: Pool the first layer globally averaged Through linear layer dimensional projection and GeLU activation function, nonlinear intermediate quantities are introduced to fit complex channel relationships to obtain intermediate features.
[0234]
[0235] in: is the linear transformation dimension-raising parameter, is the linear transformation bias parameter.
[0236] Step 2.9: Transform the intermediate features The linear projection layer is used to reduce the dimension and restore it through Layer gets the first layer channel weight .
[0237]
[0238] in: is the dimension reduction matrix parameter of the projection layer, It is the bias parameter of the projection layer dimensionality reduction.
[0239] Step 2.10: Transform the first layer of spatial weight map , the first layer channel weight Embedding features with the initial block Output the first layer of fused features through mixed gating weights .
[0240]
[0241] in, is the gated Hadamard product, which realizes regional gated attention focus. It is an element-by-element multiplication to achieve channel dimension scaling.
[0242] Step 2.11: Calculate the first layer of dual-path residual fusion features :
[0243]
[0244] Step 2.12: Repeat steps 2.4 to 2.10 to calculate layer , , and perform feature fusion to obtain the first Layer dual-path residual fusion: This embodiment uses 12 layers, and the attention mechanism of each layer gradually extracts multi-granular features from local details (shallow layer) to global semantics (deep layer). 12 layers can cover the scale range of common defects in industrial anomaly detection (such as tiny cracks to structural deformation). Compared with deeper networks (such as 24 layers), 12 layers are more adaptable to the real-time requirements of industrial scenarios in terms of GPU memory occupancy and inference speed.
[0245] Step 3.1: From Extract features from the 4th, 8th, and 12th layer coding blocks: , . Three layers are selected: shallow local features (local texture anomalies), mid-level semantic features (component-level structural anomalies), and deep global features (detecting overall functional anomalies).
[0246] in: For step 2.12 The output features after the layer encoding block features are fused.
[0247] Step 3.2: Perform a global average pooling (GAP) operation to extract channel statistics:
[0248]
[0249] in: i,j are spatial dimension coordinates (corresponding to the position on the feature map, ranging from 1×1 to 32×32): indicating that the channel dimension is fully retained, that is, all 768 channels of each spatial position (i,j) are traversed. It is used to compress the three-dimensional feature tensor (32×32×768) into a channel description vector (768 dimensions) along the spatial dimension.
[0250] Step 3.3: Integrate intermediate features across channels through lightweight mapping layers :
[0251]
[0252] in, is the dimension-upgrading matrix parameter, expanding the channel interaction capacity, is the bias vector, GeLU is a nonlinear activation function, and smooth nonlinearity is introduced.
[0253] Step 3.4: For each layer of intermediate features Generate level attention weights:
[0254]
[0255] in: is the level attention weight matrix, It is hierarchical attention bias.
[0256] Step 3.5: Normalize the level attention weights to get the level weight normalization weights .
[0257]
[0258] in: and , realizing dynamic feature selection.
[0259] Step 3.6: Layer features And the normalized attention weight Weighted fusion to obtain fusion features To balance details and semantics.
[0260]
[0261] Step 3.7: Fusion features Align the feature dimensions through linear projection to obtain the final aligned features .
[0262]
[0263] in The alignment matrix can be learned to map high-dimensional features to low-dimensional semantic space. The final aligned features are used for downstream feature interaction, classification and localization tasks.
[0264] Step 4.1: Visual features through The convolution operation is used to reduce the dimension of features and obtain compressed features .
[0265]
[0266] Step 4.2: Compress the features Generate spatial attention aggregation weights through 3×3 convolution and normalization layers .
[0267]
[0268] Step 4.3: Compress the features and spatial attention weights Perform context aggregation to obtain a global description vector :
[0269]
[0270] in, i , j ∈ [1,32 ] is the spatial position, and each and Multiply them together so that the spatial attention weights are applied to each position of the feature map.
[0271] Step 4.4: Convert the global description vector The dimension is increased to 1024 through the fully connected layer, and the hidden layer features with nonlinear expression ability are enhanced through the GeLU activation function. , the expression is as follows:
[0272]
[0273] in: is a dimension-raising matrix, is the bias matrix.
[0274] Step 4.5: Use the split projection strategy to transform the 1024-dimensional hidden layer features Decoupled into 6 groups of independent subspaces, each group generates a 512-dimensional latent vector through an independent linear layer
[0275]
[0276] in, is the head projection weight of the i-th group, is the projection bias of the ith group, is the linear transformation parameter.
[0277] Step 4.6: Convert 6 groups of 512-dimensional vectors Concatenate into dynamic hint prefix matrix, and obtain dynamic hint with stable feature distribution through layer normalization. :
[0278]
[0279] Step 4.7: Call the open source large language model Acquire industrial semantically enhanced knowledge:
[0280]
[0281] in: For large language models The temperature parameter controls the randomness of the generation. The larger the value, the more random the generation. Select the five most confident defect descriptions from the model output. It is a sampling strategy that limits the probability distribution of the generated results, that is, only words with a cumulative probability of 0.9 are considered. The prompt is "As an industrial quality inspection expert, please list [k] types of [CLASS] physical defects that may occur during the manufacturing process, including the defect shape, location and typical size."
[0282] in: The category identifier representing the input (such as 15 categories of industrial objects such as "transistor" and "leather" in the MVTec AD dataset) is used as the semantic variable of the prompt word template to guide the large model to generate defect knowledge that is strongly related to the target category.
[0283] Step 4.8: Add the dynamic prefix , the class name is concatenated with the "flawless" semantics to get a normal prompt template .
[0284]
[0285] Step 4.9: Add the dynamic prefix , concatenate the class name with the LLM semantics to get the exception prompt template
[0286]
[0287] in For sentence concatenation operations, cross-modal dynamic prompt prefixes, category names, and semantically enhanced knowledge are concatenated.
[0288] Step 4.10: Normal prompt template With exception prompt template Feed into CLIP text encoder to get normal semantic projection Projection with exception semantics .
[0289]
[0290] in: is 512, and K is the number of semantic enhancement items in step 4.7, 5.
[0291] Step 5.1:
[0292] The aligned visual features , positive and negative text features The attention calculation triples are obtained by cross-modal semantic alignment through the normalization layer and the linear layer:
[0293]
[0294] in: for learnable visual query attention weights, is the normal text key attention weight, is the attention weight of the abnormal text key, is the normal text value attention weight, is the attention weight of the abnormal text value, It is a normalization operation to stabilize the feature distribution.
[0295] Step 5.2: Visual query vector Each spatial position ( h , w ) ∈ [1,32 ] Calculate adaptive attention weights:
[0296]
[0297] in: is the adaptive positive attention weight, is the adaptive negative attention weight, is the spatial position, is the scaling factor.
[0298] Step 5.3: Send in , Computational vision drives attention interaction of normal and abnormal text , .
[0299]
[0300]
[0301] Step 5.4: Abnormal text features after attention interaction optimization Nonlinear high-order semantic features are obtained through nonlinear mapping layers .
[0302]
[0303]
[0304] in is the weight of the high-order semantic extension layer, is the weight of the high-order semantic compression layer.
[0305] Step 5.5 Replace step 4.7 The generated defect description is projected into the semantic space through the CLIP text encoder to generate domain knowledge bias items :
[0306]
[0307] Step 5.6: Transform nonlinear high-order semantic features Biased by domain knowledge Add to avoid the noise introduced by modal interaction differences. Finally, normalize to obtain semantically complete abnormal text features :
[0308]
[0309] Step 5.7 At the same time, the aligned visual features conduct Convolutional transformation to generate visual query vectors aligned with the text modality :
[0310]
[0311] Step 5.8: Copy and expand the optimized text features along the spatial dimension to generate a key vector aligned with the visual feature space. .
[0312]
[0313] Step 5.9 will , The convolution activation gating mechanism is used to enhance the local abnormal response:
[0314]
[0315] in, Logo stitching and Channel features, generate spatial gating weights through convolution and nonlinear activation functions .
[0316] Step 5.10: Optimize the text features Divide into 4 projections :
[0317]
[0318] At the same time , Projection:
[0319] in: The projection matrix for each head.
[0320] Step 5.11 , , Send it to the four-head attention for fusion to get the multi-head fusion vector :
[0321]
[0322]
[0323] Step 5.12: Set the spatial gating weight and Perform residual weighted fusion to obtain optimized visual features :
[0324]
[0325] in, It is the gated Hadamard product to achieve regional attention focus. This method can realize spatial weighting of attention results, suppress non-salient areas, and strengthen the response of abnormal areas. Gated weighted additional residual connection (connecting the original visual features ) further strengthened the response to abnormal areas.
[0326] Step 6.1: Visual features , normal text features , abnormal text features Feed in cross-modal contrast loss: Drive cross-modal feature alignment by comparing the feature similarities of image-normal text pairs and image-abnormal text pairs.
[0327]
[0328] in is the cosine similarity, is the temperature coefficient, controlling the distribution smoothness, The normal sample size is 4. The batch size is 16. By comparing the similarity between image features and normal / abnormal text features, normal image features are forced to align with normal text features while staying away from abnormal text features, thus enhancing the cross-modal feature distinguishability.
[0329] Step 6.2: Optimize the visual features , text features Input feature distribution alignment loss function:
[0330]
[0331] in: is a set of randomly sampled spatial locations ( ). It is a learnable MLP projection head that constrains the distribution consistency of interactive visual features and interactive text features to maintain the stability of cross-modal interaction.
[0332] Step 6.3: Normal text features With abnormal text features Feed in an explicit text boundary constraint loss:
[0333]
[0334] in: The minimum interval threshold determined by experimental tuning , force a minimum spacing between normal text and abnormal text features , to avoid semantic confusion.
[0335] Step 6.4: The total loss is :
[0336]
[0337] in: =1.0 is the boundary loss weight, strengthening the separation of normal and abnormal features, = 0.5 is the alignment loss weight to ensure the semantic consistency of vision and text.
[0338] Step 6.5: Define optimizer and training loop: Initialize StableAdamW optimizer and set learning rate lr to 1×e −3 , weight decay wd is 1×e −4 , and continued for 200 epochs.
[0339] Step 6.6: Forward propagation: Execute steps 2-5 in sequence to generate visual feature image features , normal text features , Abnormal text features And optimized visual features With optimized text features .
[0340] Step 6.7: Back propagation and model saving: Substitute the forward propagation features of step 6.6 into the total loss , and back-propagate to update the parameters of the industrial anomaly detection model, completing 200 training batch iterations in turn. Save the parameters of the trained dual-path visual encoder, multi-level feature adaptive fusion adapter, cross-modal dynamic prompt embedder, and cross-modal interaction module for loading in the test phase.
[0341] Step 7.0: Load the trained industrial anomaly detection model, freeze all trainable parameters, and switch to inference mode. Load the test set data and execute steps 2-5 to obtain the optimized multi-level visual features. With optimized text features .
[0342] Step 7.1: Optimize the cross-modal visual features Expand along the spatial dimension to obtain spatial features :
[0343]
[0344] Step 7.2: Calculate the cosine similarity between the visual features and the text anomaly features and perform spatial reshaping to generate the initial anomaly response map :
[0345]
[0346]
[0347] Step 7.3: Map the abnormal response pass Bilinear upsampling to 512×512 and average pooling With layer normalization Get dynamic learnable weight parameters .
[0348]
[0349] Step 7.4: Map the initial abnormal response to With weight parameters Perform weighted fusion to obtain the final abnormal response map .
[0350]
[0351] Step 7.5: Calculate image-level anomaly score:
[0352] ,in, represents the global score, Indicates local rating.
[0353] in: Combining the maximum pooling value of global similarity and local similarity, is the cosine similarity.
[0354] On the MVTec AD validation set, let the image-level anomaly score >0.85 is considered abnormal.
[0355] Step 7.6: Feed Perform pixel-level positioning:
[0356]
[0357] Where: Threshold ( is the mean, is the standard deviation) to generate anomaly masks and achieve pixel-level positioning.
[0358] Embodiment 3:
[0359] This embodiment introduces a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, a few-sample industrial anomaly detection method based on cross-modal adaptive interaction as described in any one of Embodiments 1 is implemented.
[0360] Embodiment 4:
[0361] This embodiment introduces a computer device, including:
[0362] Memory, used to store instructions.
[0363] A processor is used to execute the instructions so that the computer device performs the operations of a few-sample industrial anomaly detection method based on cross-modal adaptive interaction as described in any one of Embodiment 1.
[0364] Embodiment 5:
[0365] This example introduces an experiment conducted on the authoritative benchmark dataset MVTec AD. The core research goal of this dataset is to train an industrial anomaly detection model based on normal samples to accurately identify and classify abnormal states in test samples. The experimental method adopts a class-by-class training and testing strategy, that is, for each category, 4 normal sample images are randomly selected from the training set for few-sample model training, and the test is performed immediately after training to complete the detection tasks of all 15 categories in turn. The StableAdamW optimizer is used, and the learning rate is 1×e −3 , the weight decay is 1×e −4 , lasting 200 epochs and running on an NVIDIA GeForce RTX 2080ti GPU. The evaluation indicators are image-level AUROC (Area Under the Receiver Operating Characteristic Curve), which measures the model's ability to distinguish between normal and abnormal images; pixel-level AUROC, which evaluates the accuracy of abnormality location. As shown in Table 1:
[0366] Table 1 is a comparison table of MVTec AD ablation experiments under 4 samples
[0367]
[0368] Through the above process, this embodiment achieves image-level AUROC 96.4% and pixel-level AUROC 96.9% on the MVTec AD dataset, verifying the efficiency and practicality of this method in industrial anomaly detection.
[0369] The method of the present invention designs dual-path block coding and multi-scale fusion for high-resolution images (512×512) to balance computational efficiency and detail retention; only through a small number of sample training (4 normal samples) and multi-modal dynamic prompt generation (most of the previous methods only had visual unimodality or multi-modality that only fused static text prompts), it reduces the dependence on industrial manually labeled data; combined with LLM domain knowledge injection and cross-modal two-way interaction mechanism, it adapts to the semantic requirements of diverse working conditions and improves the detection robustness of complex defects.
[0370] Among them, the modules of the industrial anomaly detection model, such as Figure 2 As shown, it has the following advantages:
[0371] Dual-path visual encoder: Improved CLIP model, integrating global QKV attention and local spatial-channel hybrid attention (SCMA), taking into account both global semantic understanding of the image and focusing on local defect areas, solving the problem of subtle anomaly detection in complex backgrounds.
[0372] Multi-level feature adaptive fusion adapter: Dynamically weighted fusion of visual features at different levels in the encoder (shallow details, mid-level transitions, deep semantics), balancing multi-scale information through a hierarchical attention mechanism to improve the robustness of feature expression.
[0373] Cross-modal dynamic prompt embedder: Generates dynamic text prompt prefixes based on visual feature context, calls the large language model (LLM) to inject domain knowledge, builds refined normal / abnormal text templates, and enhances semantic adaptation capabilities.
[0374] Cross-modal interaction module: Through vision → text semantic modulation and text → vision gating mechanism, it realizes two-way feature optimization interaction and strengthens the response and semantic consistency of abnormal areas.
[0375] Multi-objective joint optimization loss function: Combine cross-modal contrast loss, feature distribution alignment loss and explicit text boundary constraint loss to jointly optimize model parameters to ensure feature discriminability under few-sample conditions.
[0376] Multi-scale anomaly decision and scoring: Fusion of multi-level feature similarity graphs, combined with global and local scoring, to achieve image-level anomaly determination and pixel-level defect location.
[0377] The present invention improves the structure of the visual encoder and introduces a global and local dual-path encoder to enhance the ability to focus on subtle abnormal areas while retaining the global semantic association of the image; designs a multi-level feature adaptive fusion adapter to dynamically balance the details and semantic information of visual features at different levels, and constructs a cross-modal instance adaptive prompt embedding technology to align visual features with dynamically generated text prompts to improve the accuracy of modal interaction. On this basis, combined with the domain knowledge injection strategy of the large language model (LLM), a refined and scene-adaptive dynamic text prompt template is generated to replace the traditional static template and effectively adapt to the semantic requirements of diversified working conditions. Through the joint optimization strategy of cross-modal contrast learning, feature distribution alignment and explicit text boundary constraints, the present invention can significantly enhance the model's ability to distinguish normal and abnormal features under the condition of only a small number of normal samples, reduce the dependence on labeled data, and provide a high-precision, low-cost and rapidly deployable industrial anomaly detection solution for intelligent manufacturing, promoting the industrial application of few-sample learning and cross-modal technology.
[0378] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A few-sample industrial anomaly detection method based on cross-modal adaptive interaction, characterized by: Specifically include: Get a few-sample dataset; Input the few-shot dataset into the dual-path visual encoder and get Layer dual-path residual fusion features; Will The dual-path residual fusion features of multiple layers in the layer dual-path residual fusion feature are input into the multi-level feature adaptive fusion adapter to obtain the final aligned features; The aligned final features are input into the cross-modal dynamic cue embedder to obtain normal semantic projection and abnormal semantic projection; The aligned final features, normal semantic projections, and abnormal semantic projections are input into the cross-modal interaction module to obtain optimized text features and optimized visual features. The total loss is calculated based on the aligned final features, normal semantic projection, abnormal semantic projection, optimized visual features, and optimized text features, and the total loss is used to update the parameters of the industrial anomaly detection model to obtain a trained industrial anomaly detection model. The industrial image to be detected is input into the trained industrial anomaly detection model to obtain the aligned final features, anomaly semantic projection, optimized visual features and optimized text features, which are used to judge the anomaly of the industrial image to be detected.
2. The method for detecting industrial anomalies with a small number of samples based on cross-modal adaptive interaction according to claim 1 is characterized by: Also includes: The abnormal situation is located according to the aligned final features of the industrial image to be detected, the normal semantic projection and the abnormal semantic projection, the optimized visual features and the optimized text features.
3. A method for detecting industrial anomalies with a small number of samples based on cross-modal adaptive interaction according to claim 1 or 2, characterized in that: The few-sample dataset is K normal samples randomly selected from each type of dataset as the few-sample training set. , , For normal samples.
4. A method for detecting industrial anomalies with a small number of samples based on cross-modal adaptive interaction according to claim 1 or 2, characterized in that: The few-sample dataset is input into the dual-path visual encoder to obtain Layer dual-path residual fusion features, specifically including: Step 2.1: Concentrate the minority sample data into normal samples Convolution blocks are performed sequentially to obtain block features ; Step 2.2: Sequentially block features in the channel dimension Perform layer normalization to obtain the initial block embedding features ; Step 2.3: Embedding the initial block into features Input to the global context attention path, generate the query vector for the first layer of attention calculation through three sets of independent linear projections , key vector Sum value vector ; Step 2.4: Calculate the first-level attention features , the expression is as follows: ; in, For normalization operation, is the scaling factor; Step 2.5: The first-level attention features Input feedforward network to expand feature dimension to obtain intermediate layer features , first layer QKV attention , the expression is as follows: ; ; in, is the weight parameter of the feedforward network extension layer, is the bias parameter of the feedforward network extension layer, is the weight parameter of the compression layer of the feedforward network, is the bias parameter of the compression layer of the feedforward network, is the activation function; Step 2.6: Embedding the initial block into features Input spatial-channel hybrid attention path, generate the first layer spatial weight map through lightweight single channel convolution ; Step 2.7: Calculate initial block embedding features The first layer global average pooling value ; Step 2.8: Pool the first layer globally averaged Through the linear layer dimension projection and activation function, the intermediate features are obtained ; Step 2.9: Transform the intermediate features The first layer channel weights are obtained through dimensionality reduction and restoration through the linear projection layer and activation function ; Step 2.10: Transform the first layer of spatial weight map , the first layer channel weight Embedding features with the initial block Output the first layer of fused features through mixed gating weights ; Step 2.11: Based on the first layer of QKV attention , the features after the first layer fusion and the initial block embedding features Calculate the first layer of dual-path residual fusion features , the expression is as follows: ; Step 2.12: Repeat steps 2.4 to 2.11 to calculate of Layer dual-path residual fusion feature , the expression is as follows: 。 5. The method for detecting industrial anomalies with a small number of samples based on cross-modal adaptive interaction according to claim 1 or 2, characterized in that: The The dual-path residual fusion features of multiple layers in the layer dual-path residual fusion feature are input into the multi-level feature adaptive fusion adapter to obtain the final features after alignment, including: Step 3.1: From Layer dual-path residual fusion feature Extract the , , Dual-path residual fusion features of the layer , As a hierarchical feature ,in, , , ; Step 3.2: Layer features Perform global average pooling respectively to obtain the average pooling value ,in, , , ; Step 3.3: Average pooling value Obtain intermediate features through lightweight mapping layers , the expression is as follows: ; in, is the dimension-raising matrix parameter, is the bias vector, GeLU is the activation function; Step 3.4: Calculate the intermediate features of each layer The level attention weights , the expression is as follows: ; in: is the level attention weight matrix, It is hierarchical attention bias; Step 3.5: Set the level attention weights Normalize the weights to get the hierarchical weight normalization weights ; Step 3.6: Layer features And the normalized attention weight Weighted fusion to obtain fusion features ; ; Step 3.7: Fusion features Align the feature dimensions through linear projection to obtain the final aligned features ; ; in: is the learnable alignment matrix.
6. A method for industrial anomaly detection based on cross-modal adaptive interaction with few samples according to claim 1 or 2, characterized in that: The final aligned features are input into the cross-modal dynamic prompt embedder to obtain normal semantic projection and abnormal semantic projection, specifically including: Step 4.1: Final features after alignment Obtaining compressed features through convolution operations ; Step 4.2: Compress the features Generate spatial attention aggregation weights through convolution and normalization layers ; Step 4.3: Suppress the features and spatial attention aggregation weights Perform context aggregation to obtain a global description vector ; Step 4.4: Convert the global description vector The dimension is increased through the fully connected layer, and the hidden layer features are obtained through the activation function ; Step 4.5: Use the split projection strategy to transform the hidden layer features Decoupled into S groups of independent subspaces, each group generates a hidden vector through an independent linear layer ,in, , the expression is as follows: ; in, is the head projection weight of the i-th group, is the projection bias of the ith group, is the linear transformation parameter; Step 4.6: Set S as the latent vector Concatenate into dynamic hint prefix matrix and get dynamic hint by layer normalization ; Step 4.7: Call the large language model to obtain industrial semantic enhancement knowledge , the expression is as follows: ; in: For large language models, To select the Z most confident defect descriptions from the output of the large language model, Represents the category identifier of the input; Step 4.8: Add the dynamic prefix , the class name is concatenated with the "flawless" semantics to get a normal prompt template ; ; Step 4.9: Add the dynamic prefix , class names and industrial semantics enhanced knowledge Perform splicing to get the abnormal prompt template ; ; in: It is a statement concatenation operation; Step 4.10: Normal prompt template With exception prompt template Feed into the text encoder to get normal semantic projection Projection with exception semantics .
7. The method for detecting industrial anomalies with a small number of samples based on cross-modal adaptive interaction according to claim 1 or 2, characterized in that: The aligned final features, normal semantic projections and abnormal semantic projections are input into the cross-modal interaction module to obtain optimized text features and optimized visual features, specifically including: Step 5.1: Order As positive and negative text features, the final features after alignment , positive and negative text features The visual query attention vector of attention calculation is obtained by cross-modal semantic alignment through normalization layer and linear layer respectively. , normal text key attention vector , abnormal text key attention vector , normal text value attention vector and the abnormal text value attention vector , the expression is as follows: ; in: for learnable visual query attention weights, is the normal text key attention weight, is the attention weight of the abnormal text key, is the normal text value attention weight, is the attention weight of the abnormal text value, It is a normalization operation; Step 5.2: Attention vector for visual query Calculate the adaptive positive attention weight for each spatial position of , Adaptive negative attention weight , the expression is as follows: ; in: is the spatial position, is the scaling factor; Step 5.3: Adaptive positive attention weights , Adaptive negative attention weight , calculate the normal text features after attention interaction optimization , abnormal text features after attention interaction optimization , the expression is as follows: ; ; in, represents the total value of height, and b represents the total value of width; Step 5.4: Abnormal text features after attention interaction optimization Nonlinear high-order semantic features are obtained through nonlinear mapping layers , the expression is as follows: ; ; in: is the weight of the high-order semantic extension layer, Biasing the layer for higher-level semantic extension; is the weight of the high-order semantic compression layer, High-level semantic compression layer bias; Step 5.5: Enhance knowledge with industry semantics The generated defect description is projected into the semantic space through the text encoder to generate domain knowledge bias items ; Step 5.6: Nonlinear high-order semantic features Bias with domain knowledge After addition, normalization is performed to obtain semantically complete abnormal text features. ; Step 5.7: Align the final features Perform convolution transformation to generate a visual query vector aligned with the text modality '; ; Step 5.8: Copy and expand the optimized text features along the spatial dimension to generate a key vector aligned with the visual feature space ; ; in, Represents the optimized text features, let ; Step 5.9: Align the visual query vector with the text modality , a key vector aligned with the visual feature space Generate spatial gating weights through convolution and nonlinear activation functions ; Step 5.10: Divide the H-head projections to obtain multi-head text value vectors ,Will , Projecting multi-head visual query vector , multi-headed text key vector , the expression is as follows: ; ; ; in: is the query projection matrix for each head, is the key projection matrix for each head, The value projection matrix for each head; Step 5.11: , , Input H head attention to fuse and get multi-head fusion vector , the expression is as follows: ; ; in, Output fusion weights for multiple heads; Step 5.12: Set the spatial gating weights Fusion vector with multiple heads And the final features after alignment Perform residual weighted fusion to obtain optimized visual features , the expression is as follows: ; in, is the gated product.
8. The method for detecting industrial anomalies with a small number of samples based on cross-modal adaptive interaction according to claim 1 or 2, characterized in that: The total loss is calculated based on the aligned final features, normal semantic projections, abnormal semantic projections, optimized visual features, and optimized text features, and the total loss is used to update the parameters of the industrial anomaly detection model to obtain a trained industrial anomaly detection model, specifically including: Step 6.1: Initialize the optimizer, set the learning rate, weight decay, and the number of iterations, execute steps 2-5 for forward propagation, and generate the final features after alignment , normal semantic projection Projection with exception semantics , optimized visual features And the optimized text features ; Step 6.2: Based on the final features after alignment , normal semantic projection Projection with exception semantics Compute cross-modal contrast loss , the expression is as follows: ; in, is the cosine similarity, is the temperature coefficient, is the normal sample size, For batches, , , They represent the final features, normal semantic projection and abnormal semantic projection of the ith normal sample respectively, and e is a natural constant; Step 6.3: Based on the optimized visual features , optimized text features Calculate the feature distribution alignment loss function , the expression is as follows: ; in: is a set of randomly sampled spatial locations, Align the projection head for vision Align the drop shadow header for the text, Represents the feature vector of the optimized visual feature at the spatial position (h, w), represents the norm of 2; Step 6.4: Projection based on normal semantics Projection with exception semantics Compute explicit text boundary constraint loss , the expression is as follows: ; in: is the minimum interval threshold, represents the maximum value function; Step 6.5: Based on cross-modal contrast loss , feature distribution alignment loss function and explicit text boundary constraint loss Calculate total loss , the expression is as follows: ; in: is the boundary loss weight, is the alignment loss weight; Step 6.6: Back-propagation updates the parameters of the industrial anomaly detection model, completes the number of training iterations in turn, saves the parameters of the trained dual-path visual encoder, multi-level feature adaptive fusion adapter, cross-modal dynamic prompt embedder and cross-modal interaction module, and obtains the trained industrial anomaly detection model.
9. The method for detecting industrial anomalies with a small number of samples based on cross-modal adaptive interaction according to claim 1 or 2, characterized in that: The industrial image to be detected is input into the trained industrial anomaly detection model to obtain the aligned final features, abnormal semantic projection, optimized visual features and optimized text features, which are used to judge the abnormality of the industrial image to be detected, specifically including: Step 7.1: Input the industrial image to be detected into the trained industrial anomaly detection model to obtain the final features after alignment , Abnormal semantic projection , optimized visual features And the optimized text features ; Step 7.2: Optimize the visual features Expand along the spatial dimension to obtain spatial features ; Step 7.3: Calculate spatial features Projection with exception semantics The cosine similarity of and spatial reshaping is performed to generate the initial abnormal response map , the expression is as follows: ; ; Step 7.4: Map the initial abnormal response to Through upsampling, average pooling and layer normalization, dynamic learnable weight parameters are obtained ; Step 7.5: Map the initial abnormal response to With weight parameters Perform weighted fusion to obtain the final abnormal response map ; Step 7.6: Based on the final features after alignment , optimized text features Computing image-level anomaly scores , the expression is as follows: ; in: Represents the maximum pooling value parameters of global similarity and local similarity, is the cosine similarity; Step 7.7: When image-level anomaly scoring When it is greater than the threshold, it is determined to be an abnormal situation.
10. The method for industrial anomaly detection based on cross-modal adaptive interaction with a small number of samples according to claim 2, characterized in that: The method of locating the abnormal situation according to the aligned final features of the industrial image to be detected, the normal semantic projection and the abnormal semantic projection, the optimized visual features and the optimized text features specifically includes: Step 8.1: Optimize the visual features Expand along the spatial dimension to obtain spatial features ; Step 8.2: Calculate spatial features Projection with exception semantics The cosine similarity of and spatial reshaping is performed to generate the initial abnormal response map , the expression is as follows: ; ; Step 8.3: Map the initial abnormal response to Through upsampling, average pooling and layer normalization, dynamic learnable weight parameters are obtained ; Step 8.4: Map the initial abnormal response to With weight parameters Perform weighted fusion to obtain the final abnormal response map ; Step 8.5: Based on the final abnormal response graph , calculate the anomaly mask , get the abnormal situation location; ; in, is the threshold value, and 1 is the pixel coordinate of the abnormal situation.
Citation Information
Patent Citations
Industrial anomaly detection method and system based on fine-grained text prompt feature engineering
CN118568650A
Zero sample image anomaly detection method and device
CN119130931A
Small sample anomaly detection and classification framework based on reconstruction guide cross-modal alignment
CN119762847A
Cited By
Industrial image anomaly detection method based on noise suppression modal fusion alignment
CN120298399A
Training method of behavior decision model and adaptive interaction method of digital human
CN120354176A
Printed circuit board defect detection method and system based on cross-modal prompt learning and visual guidance
CN120471929A
A printed circuit board defect detection method and system based on cross-modal prompt learning and visual guidance
CN120471929B
Small-sample defect identification method based on cross-modal text semantic driving
CN120580702A