Small-sample defect identification method based on cross-modal text semantic driving
Through the cross-modal text semantic-driven method, abnormal feature space and synthetic defect sample training models are generated, which solves the problem of insufficient generalization of the model in industrial quality inspection scenarios, and realizes accurate defect recognition and pixel-level positioning under the condition of few samples.
Patent Information
- Application Number
- CN202511089295.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-08-05
AI Technical Summary
The existing technology lacks flexible adaptability to dynamic defect types in industrial quality inspection scenarios. Traditional models require a large amount of labeled data training and are difficult to adapt to unknown scenarios, resulting in insufficient generalization of the model and attenuation of detection performance.
Through a cross-modal text semantic-driven method, normal images and text descriptions are used to generate abnormal feature spaces, an exception determination mechanism without labeling data is built, and a lightweight model is trained in combination with dynamically generated synthetic defect samples to achieve accurate detection under the condition of few samples.
Under the condition of no real abnormal data, unknown defects can be accurately identified, model iteration costs and labeling workload can be reduced, and generalization performance and capture subtle defects in complex scenarios can be improved.
Smart Images

Figure CN120580702A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image processing technology, and in particular to a few-sample defect recognition method driven by cross-modal text semantics. Background Art
[0002] Image anomaly detection is the process of identifying unusual areas or defects in an image that deviate from normal patterns. It has widespread applications in various fields, such as detecting product defects in industrial production, identifying disease symptoms in medical imaging, and monitoring machine behavior to warn of failures. In industrial quality inspection scenarios, this technology can accurately locate defects such as cracks and coating loss on part surfaces by analyzing product surface texture and structural features in real time.
[0003] Traditional methods for training anomaly detection models typically involve training a classifier model on a labeled dataset containing both normal and abnormal images. On the one hand, the dynamic changes in defect types in industrial scenarios make labeled datasets less up-to-date. On the other hand, a large amount of labeled data is required to effectively train the model. Due to the dynamic and evolving nature of anomalies, anomaly detection models can quickly become outdated. Traditional learning models often require extensive retraining and updates to adapt to these changes, which is both time-consuming and labor-intensive. In contrast, few-shot learning models are designed to learn new patterns from minimal or no labeled examples, can adapt to new contexts and situations more quickly, and require significantly less cost and effort.
[0004] In the task of few-shot anomaly detection, the fundamental lack of data for the target scene is a core challenge. Due to the unpredictability of actual scenarios, models trained directly on public data are often difficult to adapt to specific detection needs. In conventional methods, new, unoptimized models tend to overfit common data features, resulting in severe performance degradation in unknown target domains. This phenomenon stems from the model's lack of implicit knowledge of the target scene and its inability to extract cross-scenario discrimination rules from the underlying data. Therefore, building a pre-training architecture with strong transfer capabilities has become a key breakthrough. The core contradiction that current technology urgently needs to resolve is: how to guide the model to establish a universal anomaly representation through public data resources while maintaining flexible adaptability to unknown scenarios, ultimately achieving accurate detection and positioning with the support of few target domain data. Summary of the Invention
[0005] Purpose of the invention: In response to the above problems, the purpose of the present invention is to provide a few-sample defect recognition method driven by cross-modal text semantics. By parsing the intrinsic relationship between text descriptions and visual features, an anomaly judgment mechanism that does not require labeled data is established, which supports rapid model generation under few-sample conditions, significantly reduces data dependence, and improves the generalization ability of unknown anomaly types in dynamic scenes.
[0006] Technical solution: The present invention is based on a few-sample defect recognition method driven by cross-modal text semantics, which includes a training phase and a testing phase; During the training phase: Obtain global features and local features of normal images, and store the local features of normal images in a reference library; Construct normal prefixes and abnormal guide words, generate a normal description set based on the normal prefixes, generate an abnormal description set based on the abnormal guide words, and perform feature extraction to obtain a normal description feature set and an abnormal description feature set; Combine the global features of the normal image with the normal description features and the abnormal description features to generate a normal feature vector and an abnormal feature vector; Using the abnormal feature vector to obtain a synthetic abnormal image; Label the normal images and the synthetic abnormal images, use the normal images and the synthetic abnormal images to form a training set, and use the training set to train the binary classification model; During the testing phase: Use normal images and defect images to form a test set; Select any image in the test set as the image to be detected, and extract the global image features and local image features of the image to be detected; The distance between the normal feature vector and the abnormal feature vector and the global image features of the image to be detected is calculated using the nearest neighbor algorithm, and the image-level abnormality score is generated based on the distance difference; Input the image to be detected into the trained binary classification model to obtain the prediction result; The image-level anomaly score and the prediction result are combined to obtain the image-level anomaly detection score, which is then compared with the preset threshold to obtain the judgment result.
[0007] Furthermore, the step of obtaining the global features and local features of the normal image includes: The preprocessed normal image is input into the image encoder, where it is divided into pixel block sequences by the block embedding layer and extracted by the multi-head self-attention mechanism to obtain the global features. and local features , where c represents the number of image blocks after the spatial image is divided, d represents the feature dimension of each image block encoded in the visual language model Clip, and the local feature map of the normal image is stored in the reference library S. The local feature map of the normal image in the reference library is used to calculate the cosine similarity with the local feature map of the query image in the test phase.
[0008] Furthermore, the steps of generating a normal description set based on the normal prefix and generating an abnormal description set based on the abnormal guide word include: Let the normal prefix be , Indicates the total number of normal prefixes that can be evolved. Indicates one of the evolveable normal prefixes; Normal prefix prompt word The target object name is fused with the bidirectional attention mechanism and constructed as follows: Normal semantic templates form N normal description templates, where Indicates the target object category name, attention mechanism Using the scaled dot product form, Represents the concatenation operation of a vector sequence; The abnormal guide word is recorded as , Indicates the total number of abnormal guide words, Indicates one of the abnormal guide words; Change the category name Spliced after each abnormal guide word, a set of specific abnormal text descriptions is obtained, expressed as , the template exception description set is ,form abnormal semantic templates, where M represents the total number of categories of corresponding labeled defects; Through the dynamic weight fusion module Abnormal guide word sequence With normal semantic template The association is constructed in the form of The evolvable abnormal template forms Evolvable anomaly templates, among which, represents bilinear feature interaction, Control the size of evolvable templates for hyperparameters; Will include A collection of normal description templates { } as a normal description set, exception semantic templates and A collection of evolvable exception templates As a collection of exception descriptions.
[0009] Furthermore, the steps of performing feature extraction to obtain a normal description feature set and an abnormal description feature set include: Normal description set { } is input into the text encoder to obtain a normal description feature set, which is expressed as ;in A text encoder representing the visual language model Clip; Set exception description Input into the text encoder to obtain the abnormal description feature set, which is expressed as: .
[0010] Furthermore, the normal description feature set and the abnormal description feature set are optimized, and the process includes: Global Features , normal description feature set and anomaly description feature set Input to the control boundary module to calculate the mean of the normal descriptive characteristics. The formula is: , Where, Indicates the normal description template index; Construct the feature similarity loss function, the formula is: , Where, represents the natural exponential function, is the temperature coefficient, represents the negative normalized dot product similarity, Represents a feature vector in the template-type exception description feature, represents a feature vector in the evolvable anomaly template; the loss Enforce global features of normal images With normal descriptive characteristics mean The similarity is higher than its similarity to the abnormal description features; Implement multi-prototype space constraints and calculate the mean of anomaly description features. The formula is: , Calculate the center of the artificial anomaly template using the formula: , Calculate the center of the evolvable abnormal template using the formula: , Construct a dual distance loss function, the formula is: , in, Represents global features The expected value of represents the square of the Euclidean distance, is the distance boundary threshold, To distribute the alignment strength coefficient, the first term constrains the global features of the normal image distance The Euclidean distance of Nearly at least Unit, the second term narrows the distribution difference between the two types of abnormal prototypes; The overall optimization goal is , after updating the evolvable parameters through gradient descent, the optimized normal description set {}, and by exception semantic templates and An exception description set consisting of an evolvable exception template , the normal description feature mean and the mean of the anomaly descriptive features Stored in the reference library S.
[0011] Furthermore, the steps of combining the global features of the normal image with the normal description features and the abnormal description features to generate a normal feature vector and an abnormal feature vector include: Global Features Through the linear projection layer Mapped to , so that the global features are aligned with the normal description features and abnormal description feature dimensions; For each normal description feature Performing cross-modal fusion: using a gating mechanism Generate normal eigenvectors , the formula is: , Where, represents element-wise product, is the evolvable parameter matrix, Sigmoid activation; Describe the characteristics of each anomaly Calculating cross-modal attention gating weights , is an evolvable parameter; hybrid features are generated by gating weights , all After concatenation along the channel dimension, it passes through the fully connected layer Projected as the final abnormal feature vector , the formula is: ; Store normal eigenvectors and abnormal feature vector To the reference library .
[0012] Furthermore, the step of obtaining a synthetic abnormal image using the abnormal feature vector includes: Using abnormal feature vectors Constructing a set of abnormal feature vectors ,Will Input into the decoding module of the pre-trained text-image generator to obtain a tensor , represents a pre-trained image-text generator that generates images based on text features. Represents the number of normal samples in the training data set, H represents the height of the synthetic abnormal image finally output by the generator, and W represents the width of the synthetic abnormal image finally output by the generator; When the network reaches the final layer, the pixel value range of the tensor z is constrained by the hyperbolic tangent function, and the original output is compressed to The interval is expressed as: , Where, It is represented as the output tensor after compression by the hyperbolic tangent function, that is, the pixel value of the synthesized image; Then, the values are mapped to the [0,255] range of the standard RGB image through linear transformation to obtain the detection samples containing the synthetic abnormal image. , the formula is: , Through the constraints Achieve resolution from Gradually improve to ,in The initial feature map base size parameter for the input generation network, For the The upsampling ratio of the layer deconvolution, is the network depth; For each sample and each exception description ] perform the generation operation independently, and the final generated set of synthetic abnormal images is expressed as: , It represents the synthetic abnormal image generated by combining the nth normal prompt template corresponding to the kth normal image sample and the mth abnormal description.
[0013] Furthermore, in the testing phase, the steps of extracting global image features and local image features of the image to be detected include: The normalized RGB query image is denoted as , where the resolution is fixed at H=W=256, Input image encoder, extract features through the multi-head self-attention mechanism of 12-layer Transformer block to obtain global image features , the formula is: , Where, represents a nonlinear activation function, is the original global semantic vector output by the image encoder, is the evolvable orthogonal projection matrix, Representation layer normalization; Will Input the multi-scale feature extractor, which consists of a 4-level convolution block, and obtains four sets of local feature maps, namely: , , , , Where, 、 、 、 They represent the first-level convolution block, the second-level convolution block, the third-level convolution block, and the fourth-level convolution block respectively. Each convolution block contains three 3×3 convolution layers and ReLU activations. For each level of local feature map Perform spatial dimension expansion and L2 normalization to generate a local image feature set. The formula is: , Where, , Indicates the Level feature map in spatial position The eigenvector at represents the Euclidean norm of a vector; In order to align with the local feature dimensions used in the training phase, the local image feature set Through the linear mapping layer Projecting to the feature space with the same dimension and semantic space as the text features extracted by CLIP, the final local image feature set is obtained = , where each local eigenvector is normalized on the unit Euclidean sphere.
[0014] Furthermore, the steps of calculating the distance between the normal feature vector and the abnormal feature vector and the global image features of the image to be detected by using the nearest neighbor algorithm, and generating the image-level anomaly score according to the distance difference include: Calculate global features separately and store in reference library Normal eigenvectors in , abnormal feature vector The cosine similarity between them is used to evaluate the similarity between the test image and the training features. The formulas are: , ; Where, Represents the L2 norm of the vector; Calculate the image-level anomaly score using the formula: .
[0015] Furthermore, the steps of obtaining an image-level anomaly detection score by integrating the image-level anomaly score and the prediction result, and comparing the image-level anomaly score with a preset threshold to obtain a determination result include: Absolute value weighted fusion and classification probability The fractional operation is used to obtain the image-level anomaly detection score , the formula is: , Where, represents the fusion weight coefficient, Represents the binary classification prediction result; When the image-level anomaly score When the image to be detected is determined to be an abnormal image, represents the classification decision threshold determined by receiver operating characteristic curve analysis.
[0016] Furthermore, when the image to be detected is an abnormal image, the predicted abnormal segmentation map of the abnormal image is detected, and the process includes: For each spatial position in the local feature map P The eigenvector of , where the eigenvector of each spatial position is Corresponding to a local area of the image, calculate The cosine similarity with the mean of normal description features and the mean of abnormal description features is as follows: , Where, is the temperature parameter, represents the vector inner product; Represents the normalized local features represents the mean of the normal descriptive features in the reference library S, represents the mean value of the abnormal description features in the reference library S; pass Constructing a preliminary semantic anomaly score map , each element in the semantic anomaly score graph ,The larger the value, the higher the probability of regional anomaly; And select the reference library S The minimum similarity value of feature vectors at the same level , the formula is: , Where, express A local visual feature vector in Indicates the local features of the test image currently being processed The feature map level; The minimum similarity value Mapped to anomaly score, the formula is: , By measuring the minimum matching degree between the local features of the query image and the local features of the normal image, the similarity range [-1,1] is linearly converted to [0,1]. The larger the value, the higher the probability of abnormality. All abnormal scores are combined to form an abnormal heat map. , spatial resolution and input feature map Consistency directly reflects the degree to which each region deviates from the normal pattern; calculate and The nonlinear sum fusion number is used as the predicted anomaly score map , the formula is: , Where, Represents the predicted anomaly score map The value of each spatial location; The predicted anomaly score graph Perform multi-scale morphological optimization and sub-pixel boundary refinement operations, where the predicted anomaly score map first generates an initial binary mask through dynamic threshold segmentation: using the improved OTSU algorithm to calculate the adaptive segmentation threshold , the formula is: , Where, is the inter-class variance of the grayscale histogram of the heat map, is the local contrast compensation factor, is a sliding window, is the local gradient amplitude, is the gradient operator, and denote the horizontal and vertical partial derivatives respectively, is the total number of pixels; If and only if When the binary mask =1 indicates abnormality, otherwise =0 means normal, generate initial mask .
[0017] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: This paper effectively solves the problem of insufficient model generalization caused by the scarcity of abnormal samples and the dynamic evolution of defect types in industrial quality inspection scenarios through a cross-modal text semantic drive and dynamic feature recombination mechanism. With the help of cross-modal semantic fusion technology, the abnormal feature space can be automatically constructed with only normal images and text descriptions. Combined with dynamically generated synthetic defect samples, the lightweight model can be trained to accurately identify unknown defects even in the absence of real abnormal data, significantly reducing the cost of model iteration and the workload of annotation. By realizing the coordinated optimization of pixel-level anomaly localization and image-level judgment through multimodal feature recombination and implicit contrast learning, the present invention can effectively handle industrial scenarios such as multi-object stacking and complex backgrounds. Its adaptive boundary discrimination mechanism can dynamically adjust the detection threshold according to the regional characteristics of the object, and accurately mark the defect boundary of each pixel by fusion average. While reducing the manual re-inspection rate, it improves the ability to capture subtle defects and the generalization performance of complex scenes, and can maintain stable detection capabilities under conditions of few samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a diagram of the architecture of the present invention during the training phase; Figure 2 This is the flow chart of the training phase; Figure 3 This is a schematic diagram of the structure of the prompt generation module; Figure 4 It is a structural diagram of the control boundary module; Figure 5 It is a structural diagram of the feature synthesis module; Figure 6 This is a diagram of the architecture of the present invention during the testing phase; Figure 7 This is the flow chart of the testing phase. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of this application more clear, this application is further described in detail below with reference to the accompanying drawings and embodiments.
[0020] The few-sample defect recognition method based on cross-modal text semantics driving described in this embodiment includes two main stages: a training stage and a testing stage. It uses a pre-trained visual language model to extract a joint representation of image and semantic information, and constructs a detection framework that can dynamically adapt to different sample scenes. The present invention establishes an anomaly determination mechanism that does not require labeled data by parsing the intrinsic relationship between text descriptions and visual features, and supports rapid model generation under few-sample conditions. Compared with the traditional training model that relies on massive labeled samples, the present invention achieves lightweight modeling through semantic driving, significantly reduces data dependence, and at the same time improves the generalization ability of unknown anomaly types in dynamic scenes.
[0021] This invention does not directly use real abnormal images for training, but instead generates multi-scale feature extraction from text descriptions. During the training phase, a normal image without abnormalities (e.g., a standard image of an industrial product) and text describing normal or abnormal conditions (e.g., "defect-free" and "damaged") are first used. The normal image is fed through a pre-trained visual language model, Clip, to output global semantic features and a local spatial feature map representing local details. The local spatial feature map of the normal image is then stored in a visual reference library to establish a reference baseline. A semantic combination of the normal prefix and the target object category name is performed to generate a set of normal descriptions. Dynamic weight fusion is used to associate the abnormal text suffix with the normal description template to generate a set of abnormal descriptions containing abnormal semantics. The text encoder in the Clip model extracts features from the normal description set, converting the abnormal descriptions into abnormal description features. Next, the normal description features and the generated abnormal description features are fed into a control boundary module, where a hyperparameter is introduced to ensure that the distance between the normal global features and the normal description features is smaller than the distance between the normal global features and the abnormal description features.
[0022] The global features of normal images are mean-fused with the features of normal text to form a combined feature vector, which is used to define the comprehensive characteristics of "normal objects." At the same time, the abnormal description is combined with the normal image features to generate synthetic abnormality vectors. These vectors are input into the text-to-image generator, which outputs synthetic images with specific defect morphologies, simulating real-world defects (such as cracks and stains). Normal images are labeled "normal" and synthetic abnormal images are labeled "abnormal." The original normal images and synthetic abnormal images together constitute the training dataset used to train a binary classification machine learning model to distinguish between normal and abnormal images.
[0023] During the testing phase, a query image is input, and the Clip model's image encoder is used to extract its global image features and local spatial feature maps. The nearest neighbor algorithm is used to calculate the distance between the query image's global features and the normal and abnormal feature vectors generated during the training phase. The closest distance between the query image's global features and the normal feature vector represents "normal similarity," while the closest distance to the abnormal feature vector represents "abnormal similarity." The difference between these two distances generates an image-level anomaly score, which reflects the degree to which the query image deviates from normality. The query image is then fed into the trained binary classification model to obtain a prediction result. Finally, the image-level anomaly detection score is calculated by combining the image-level anomaly score with the prediction result.
[0024] The local spatial feature maps of normal images stored in the reference library are compared pixel by pixel with the local feature map of the query image, generating a high-resolution anomaly heatmap using the minimum distance metric. Cosine similarity is calculated between the query image's local feature map and the mean of the normal and anomaly descriptive features stored in the reference library to generate a semantic anomaly score map. The anomaly heatmap and semantic anomaly score map are fused using a nonlinear summation formula to generate a predicted anomaly score map. The predicted anomaly score map is then thresholded and binarized to produce a pixel-level anomaly mask that precisely identifies the defect location.
[0025] The CLIP model is a core module used in the training and testing phases. Its input consists of a normal image (for training) or a test image (for testing) along with positive and negative text descriptions (e.g., "normal," "damaged"). The image encoder converts the image into image features, while the text encoder converts the text description into the corresponding text feature vector. The output is image and text features in a unified embedding space, enabling cross-modal comparison of images and text, providing foundational support for subsequent feature fusion and anomaly detection.
[0026] The training dataset consists entirely of normal RGB images of objects to be detected, while the test dataset contains both normal and defective images. During the training phase, only a small number of normal image samples are required.
[0027] Combine Figure 1 and Figure 2 As shown, the training phase includes the following steps: Step 11: Obtain the global features and local feature maps of the normal image, and store the local feature maps of the normal image in a reference library.
[0028] Furthermore, the steps of obtaining the global features and local feature maps of the normal image include: The preprocessed normal image is input into the image encoder, where it is divided into pixel block sequences by the block embedding layer and extracted by the multi-head self-attention mechanism to obtain the global features. and local feature maps , where c represents the number of image blocks after the spatial image is divided, and d represents the feature dimension of each image block encoded in the visual language model Clip.
[0029] In one example, a structured positive and negative semantic template is constructed, including a positive text library and a negative text library. The positive text library integrates basic descriptions (such as "the surface is smooth and flawless") and enhanced context combinations (such as "intact {category} photographed from multiple angles"). The negative text library is generated based on a knowledge graph of defect types, covering morphological anomalies (such as "edge cracks"), texture defects (such as "corrosion spots") and complex semantic disturbances (such as "damage {category} in low-contrast images"), ensuring that the scale of positive and negative texts reaches 200 and 300 respectively.
[0030] For a set of K defect-free industrial product RGB images { The images are preprocessed to a uniform resolution of 512×512 pixels. Each frame contains 1 to 3 target objects to avoid occlusion and overlap. The images are then normalized: pixel values are normalized to [-1, 1], and a cubic polynomial interpolation method based on the Mitchell-Netravali kernel function is used. The interpolation formula is: , Where x represents the relative distance between the pixel sampling position and the adjacent integer pixel (in pixels, with direction), Represents the interpolation weight at the position with distance x. The parameter a=-0.5 is used to balance the interpolation sharpness and smoothness. This value ensures that the geometric deformation error is less than 0.5 pixels. The background noise is segmented and filtered out by the segmentation method with a fixed threshold of 0.5.
[0031] To ensure that the effective area of the target object is ≥85%, it is necessary to implement the area ratio constraint: let the area of the original segmented area be The effective area retained is , must meet , discrete noise points are removed through regional connectivity analysis to meet the standard. The center crop specification is determined based on the minimum bounding rectangle of the target area. Suppose the original image size is , the center coordinates of the target area are , the cropping window size is preset (For example, take 80% of the short side of the image as the benchmark), and the final cropping range is symmetrically extended around the center point, that is, , and is restricted to the image boundary to ensure that key features are concentrated within the cropping box. This setting balances the effective area retention rate and key information integrity through area threshold constraint and geometric center positioning.
[0032] The Clip model includes an image encoder and a text encoder. The image encoder uses the ViT-B / 16 variant. The preprocessed image is input into the image encoder and divided into a 16×16 pixel block sequence through the block embedding layer. The global semantic features are extracted through the multi-head self-attention mechanism. , in this example, the feature dimension d=768, and the local spatial feature map , where 196 corresponds to a 14×14 spatial grid. Each grid point is associated with a 768-dimensional feature vector, preserving fine-grained structural information at a scale of 1 / 16 of the original image. The local feature maps of the normal image are stored in a reference library S. The local feature maps of the normal image in the reference library are used to calculate cosine similarity with the local feature maps of the query image in the test phase.
[0033] Step 12: construct a normal prefix and an abnormal guide word, generate a normal description set based on the normal prefix, generate an abnormal description set based on the abnormal guide word, and perform feature extraction to obtain a normal description feature set and an abnormal description feature set.
[0034] The prompt generation module (PGM) inputs a normal text prefix and an abnormal guide word. The normal text prefix is concatenated with the object name to generate a normal description. The normal description is then concatenated with the abnormal guide word to generate an abnormal description containing abnormal semantics.
[0035] In an example, normal prefixes include a clean and a normal; abnormal suffixes include 'damaged {}', 'broken {}', '{} with flaw', '{} with defect', '{} with damage', etc., and the category name can be placed in {}.
[0036] Combine Figure 3 Furthermore, the steps of generating a normal description set based on the normal prefix and generating an abnormal description set based on the abnormal guide word include: Let the normal prefix be , Indicates the total number of normal prefixes that can be evolved. Indicates one of the evolveable normal prefixes; Normal prefix prompt word The target object name is fused with the bidirectional attention mechanism and constructed as follows: Normal semantic templates form N normal description templates, where Indicates the target object category name, attention mechanism Using the scaled dot product form, Represents the concatenation operation of a vector sequence; The abnormal guide word is recorded as , Indicates the total number of abnormal guide words, Indicates one of the abnormal guide words; Change the category name Spliced after each abnormal guide word, a set of specific abnormal text descriptions is obtained, expressed as The template exception description set is ,form abnormal semantic templates, where M represents the total number of categories of corresponding labeled defects; Through the dynamic weight fusion module Abnormal guide word sequence With normal semantic template The association is constructed in the form of The evolvable abnormal template forms Evolvable anomaly templates, among which, represents bilinear feature interaction, Control the size of evolvable templates for hyperparameters; Will include A collection of normal description templates { } as a normal description set, exception semantic templates and A collection of evolvable exception templates As a collection of exception descriptions.
[0037] Furthermore, the steps of performing feature extraction to obtain a normal description feature set and an abnormal description feature set include: Normal description set { } is input into the text encoder of the Clip model to obtain a normal description feature set, which is expressed as ;in A text encoder representing the visual language model Clip; Set exception description Input into the text encoder to obtain the abnormal description feature set, which is expressed as: .
[0038] Combine Figure 4 ,Further, the normal description feature set and the abnormal description feature set are optimized, and the process includes: Construct a dual-path feature constraint mechanism to transform the global features , normal description feature set and anomaly description feature set Input to the control boundary module (CMM) to calculate the mean of the normal descriptive features. The formula is: , Where, Indicates the normal description template index; Construct the feature similarity loss function, the formula is: , Where, represents the natural exponential function, is the temperature coefficient, Represents a feature vector in the template-type exception description feature, Represents a feature vector in the evolvable anomaly template, represents the negative normalized dot product similarity, is the feature dimension; the similarity loss Enforce global features of normal images With normal descriptive characteristics mean The similarity is higher than its similarity to the abnormal description features; Implement multi-prototype space constraints and calculate the mean of anomaly description features , Artificial Anomaly Template Center Evolvable Abnormal Template Center , construct the dual distance loss function, the formula is: , in, Represents global features The expected value of represents the square of the Euclidean distance, is the distance boundary threshold, To distribute the alignment strength coefficient, the first term constrains the global features of the normal image distance The Euclidean distance of Nearly at least Unit, the second term narrows the distribution difference between the two types of abnormal prototypes; The overall optimization goal is , after updating the evolvable parameters through gradient descent, the optimized normal description set { }, and by exception semantic templates and An exception description set consisting of an evolvable exception template , after optimization, the normal description set is closer to the global features of the normal image, the normal description and the abnormal description are separated more, and the mean of the normal description feature is and the mean of the anomaly descriptive features Deposit to reference library .
[0039] Step 13: Combine the global features of the normal image with the normal description features and the abnormal description features to generate a normal feature vector and an abnormal feature vector.
[0040] Furthermore, the steps of combining the global features of the normal image with the normal description features and the abnormal description features to generate a normal feature vector and an abnormal feature vector include: The global features of the normal image , normal description features and the abnormal description feature set Input into the feature synthesis module (FSM), the structural diagram of the feature synthesis module is as follows Figure 5 As shown, in the feature synthesis module, the global feature Through the linear projection layer Mapped to , so that the global features are aligned with the normal description features and abnormal description feature dimensions; For each normal description feature Performing cross-modal fusion: using a gating mechanism Generate normal eigenvectors , the formula is: , Where, represents element-wise product, is the evolvable parameter matrix, Sigmoid activation; Describe the characteristics of each anomaly Calculating cross-modal attention gating weights , is an evolvable parameter; hybrid features are generated by gating weights , all After concatenation along the channel dimension, it passes through the fully connected layer Projected as the final abnormal feature vector , the formula is: ; Store normal eigenvectors and abnormal feature vector To the reference library .
[0041] Normal description feature set Global features of normal images Perform feature alignment using bilinear fusion ,in . Linear projection layer in the module , gating parameters 、 The learning rate is 5× End-to-end optimization achieves normal-abnormal semantic decoupling and dynamic feature space partitioning.
[0042] During the training phase, the feature synthesis module receives input from the global features of normal images, normal descriptive features, and abnormal descriptive features generated by the Clip model. The global features and normal descriptive features are fused into a combined feature vector, representing the comprehensive characteristics of a "normal object." The output is used to define a reference baseline for normal conditions, enhancing the model's ability to represent normal features. The abnormal descriptive features are then concatenated with the global features of the normal image to generate a synthetic abnormality vector, which is then used to generate abnormal images.
[0043] Step 14: Obtain a synthetic abnormal image using the abnormal feature vector.
[0044] Furthermore, the step of obtaining a synthetic abnormal image using the abnormal feature vector includes: Using abnormal feature vectors Constructing a set of abnormal feature vectors ,Will Input into the decoding module of the pre-trained text-image generator to obtain a tensor , represents a pre-trained image-text generator that generates images based on text features. Represents the number of normal samples in the training data set, H represents the height of the synthetic abnormal image finally output by the generator, and W represents the width of the synthetic abnormal image finally output by the generator; When the network reaches the final layer, the pixel value range of the tensor z is constrained by the hyperbolic tangent function, and the original output is compressed to The interval is expressed as: , Where, Represented as the output tensor after compression by the hyperbolic tangent function, that is, the pixel value of the synthesized image; this operation compresses the original output to interval; Then, the values are mapped to the [0,255] range of the standard RGB image through linear transformation to obtain the detection samples containing the synthetic abnormal image. , the formula is: , Through the constraints Achieve resolution from Gradually improve to , ensuring that the resolution improvement multiple strictly matches the network depth, where The initial feature map basic size parameter corresponding to the input generation network, For the The upsampling ratio of the layer deconvolution, is the network depth. In this example, H=256 is set. =8, ; For each sample and each exception description ] perform the generation operation independently, and the final generated set of synthetic abnormal images is expressed as: , It represents the synthetic abnormal image generated by combining the nth normal prompt template corresponding to the kth normal image sample and the mth abnormal description.
[0045] Synthetic Anomaly Image Set This data will be used as training data to input into the subsequent classification model.
[0046] GAN, the image generation module during the training phase, takes as input an anomaly feature vector and textual features of the anomaly. Based on the combination of textual descriptions and image features, it generates synthetic anomaly images that mimic real defects (such as cracked industrial parts). These synthetic anomaly images, along with real normal images, form the training set, addressing the lack of real anomaly data and improving the model's generalization capabilities for abnormal scenarios.
[0047] In step 15, the normal images and the synthesized abnormal images are labeled, a training set is formed using the normal images and the synthesized abnormal images, and a binary classification model is trained using the training set.
[0048] Normal images are marked as "normal" and synthetic abnormal images are marked as "abnormal". A two-stage optimization strategy is used during training: the linear projection layer is fixed for the first 50 cycles. Parameters, only optimize the classification head parameters , the initial learning rate is set to 5×1 , after 150 cycles, all parameters are jointly optimized, and the learning rate =1×1 , the optimizer momentum parameter , weight decay coefficient The final output of the binary classification model is : .
[0049] Combine Figures 6 and 7 As shown, the testing phase includes the following steps: Step 21: Use normal images and defective images to form a test set.
[0050] In this example, an image in which the object has no abnormalities or defects is regarded as a normal image; an image in which the object has abnormalities, such as damage, is regarded as a defective image.
[0051] In step 22, a picture is randomly selected from the test set as the image to be detected, and the global image features and local image features of the image to be detected are extracted.
[0052] Furthermore, in the testing phase, the steps of extracting global image features and local image features of the image to be detected include: The normalized RGB query image is denoted as , where the resolution is fixed at H=W=256, Input image encoder, extract features through the multi-head self-attention mechanism of 12-layer Transformer block to obtain global image features , the formula is: , Where, represents a nonlinear activation function, is the original global semantic vector output by the image encoder, is an evolvable orthogonal projection matrix; Representation layer normalization; Will Input the multi-scale feature extractor, which consists of a 4-level convolution block, and obtains four sets of local feature maps, namely: , , , , Where, 、 、 、 They represent the first-level convolution block, the second-level convolution block, the third-level convolution block, and the fourth-level convolution block respectively. Each convolution block contains three 3×3 convolution layers with ReLU activation; For each level of local feature map Perform spatial dimension expansion and L2 normalization to generate a local image feature set. The formula is: , Where, , Indicates the Level feature map in spatial position The eigenvector at Represents the Euclidean norm of the vector. In order to align with the local feature dimension used in the training phase, Through the linear mapping layer Projected into a unified feature space, we finally get = , where each local eigenvector is normalized on the unit Euclidean sphere.
[0053] The global image features and local image feature set Together they constitute a bimodal feature group ( , ) is directly transferred to the anomaly scoring module, where Drive image-level anomaly classification, Used to generate pixel-level positioning heatmaps.
[0054] In step 23, the distances between the normal feature vector and the abnormal feature vector and the global image features of the image to be detected are calculated using a nearest neighbor algorithm, and an image-level abnormality score is generated based on the distance difference.
[0055] Furthermore, the steps of calculating the distance between the normal feature vector and the abnormal feature vector and the global image features of the image to be detected by using the nearest neighbor algorithm, and generating the image-level anomaly score according to the distance difference include: Calculate global features separately and store in reference library Normal eigenvectors in , abnormal feature vector The cosine similarity between them is used to evaluate the similarity between the test image and the training features. The formulas are: , ; Where, Represents the L2 norm of the vector; Calculate the image-level anomaly score using the formula: ; Step 24: input the image to be detected into the trained binary classification model to obtain a prediction result; Step 25: The image-level anomaly score and the prediction result are combined to obtain an image-level anomaly detection score, and the image-level anomaly score is compared with a preset threshold to obtain a judgment result.
[0056] Furthermore, the steps of obtaining an image-level anomaly detection score by integrating the image-level anomaly score and the prediction result, and comparing the image-level anomaly score with a preset threshold to obtain a determination result include: Absolute value weighted fusion and classification probability The fractional operation is used to obtain the image-level anomaly detection score , the formula is: , Where, represents the fusion weight coefficient, Represents the binary classification prediction result; When the image-level anomaly score When the image to be detected is determined to be an abnormal image, represents the classification decision threshold determined by receiver operating characteristic curve analysis.
[0057] Furthermore, when the image to be detected is an abnormal image, the predicted abnormal segmentation map of the abnormal image is detected, and the process includes: For each spatial position in the local feature map P The eigenvector of , where the eigenvector of each spatial position is Corresponding to a local area of the image, calculate The cosine similarity with the mean of normal description features and the mean of abnormal description features is as follows: , Where, is the temperature parameter, the default value is 0.01, represents the vector inner product; Represents the normalized local features represents the mean of the normal descriptive features in the reference library S, represents the mean value of the abnormal description features in the reference library S; pass Constructing a preliminary semantic anomaly score map , each element of which ,The larger the value, the higher the probability of abnormality in the area; And select The minimum similarity value of feature vectors at the same level , the formula is: , Where, express A local visual feature vector in Indicates the local features of the test image currently being processed The feature map level; The minimum similarity value Mapped to anomaly score, the formula is: , By measuring the minimum matching degree between the local features of the query image and the local features of the normal image, the similarity range [-1,1] is linearly converted to [0,1]. The larger the value, the higher the probability of abnormality. All abnormal scores are combined to form an abnormal heat map. , whose spatial resolution is similar to the input feature map Consistency directly reflects the degree to which each region deviates from the normal pattern; calculate and The nonlinear sum fusion number is used as the predicted anomaly score map , the formula is: , Where, Represents the value of each spatial location of the predicted anomaly score map; The predicted anomaly score graph , performs multi-scale morphological optimization and sub-pixel boundary refinement operations, where the predicted anomaly score map first generates an initial binary mask through dynamic threshold segmentation: the improved OTSU algorithm is used to calculate the adaptive segmentation threshold , the formula is: , Where, is the inter-class variance of the grayscale histogram of the heat map, is the local contrast compensation factor, is a sliding window, is the local gradient amplitude, is the gradient operator, and denote the horizontal and vertical partial derivatives respectively, is the total number of pixels; If and only if When the binary mask =1 indicates abnormality, otherwise =0 means normal, generate initial mask .
Claims
1. A cross-modal text semantic-driven few-shot defect recognition method, characterized by: It includes training phase and testing phase; During the training phase: Obtain global features and local features of normal images, and store the local features of normal images in a reference library; Construct normal prefixes and abnormal guide words, generate a normal description set based on the normal prefixes, generate an abnormal description set based on the abnormal guide words, and perform feature extraction to obtain a normal description feature set and an abnormal description feature set; Combine the global features of the normal image with the normal description features and the abnormal description features to generate a normal feature vector and an abnormal feature vector; Using the abnormal feature vector to obtain a synthetic abnormal image; Label the normal images and the synthetic abnormal images, use the normal images and the synthetic abnormal images to form a training set, and use the training set to train the binary classification model; During the testing phase: Use normal images and defect images to form a test set; Select any image in the test set as the image to be detected, and extract the global image features and local image features of the image to be detected; The distance between the normal feature vector and the abnormal feature vector and the global image features of the image to be detected is calculated using the nearest neighbor algorithm, and the image-level abnormality score is generated based on the distance difference; Input the image to be detected into the trained binary classification model to obtain the prediction result; The image-level anomaly score and the prediction result are combined to obtain the image-level anomaly detection score, which is then compared with the preset threshold to obtain the judgment result.
2. The cross-modal text semantic-driven few-shot defect recognition method according to claim 1 is characterized in that: The steps of obtaining the global features and local features of a normal image include: The preprocessed normal image is input into the image encoder, where it is divided into pixel block sequences by the block embedding layer and extracted by the multi-head self-attention mechanism to obtain the global features. and local features , where c represents the number of image blocks after the spatial image is divided, d represents the feature dimension of each image block encoded in the visual language model Clip, and the local feature map of the normal image is stored in the reference library S.
3. The cross-modal text semantic-driven few-sample defect recognition method according to claim 2 is characterized in that: The steps of generating a normal description set based on a normal prefix and generating an abnormal description set based on an abnormal guide word include: Let the normal prefix be , Indicates the total number of normal prefixes that can be evolved. Indicates one of the evolveable normal prefixes; Normal prefix prompt word The target object name is fused with the bidirectional attention mechanism and constructed as follows: Normal semantic templates form N normal description templates, where Indicates the target object category name, attention mechanism Using the scaled dot product form, Represents the concatenation operation of a vector sequence; The abnormal guide word is recorded as , Indicates the total number of abnormal guide words, Indicates one of the abnormal guide words; Change the category name Spliced after each abnormal guide word, a set of specific abnormal text descriptions is obtained, expressed as , the template exception description set is ,form abnormal semantic templates, where M represents the total number of categories of corresponding labeled defects; Through the dynamic weight fusion module Abnormal guide word sequence With normal semantic template The association is constructed in the form of Evolvable abnormal template, forming Evolvable anomaly templates, among which, represents bilinear feature interaction, Control the size of evolvable templates for hyperparameters; Will include A collection of normal description templates { } as a normal description set, exception semantic templates and A collection of evolvable exception templates As a collection of exception descriptions.
4. The cross-modal text semantic-driven few-shot defect recognition method according to claim 3 is characterized in that: The steps of extracting features to obtain a normal description feature set and an abnormal description feature set include: Normal description set { } is input into the text encoder to obtain a normal description feature set, which is expressed as ;in A text encoder representing the visual language model Clip; Set exception description Input into the text encoder to obtain the abnormal description feature set, which is expressed as: .
5. The method for identifying defects with a small number of samples based on cross-modal text semantics according to claim 4 is characterized in that: The normal description feature set and the abnormal description feature set are optimized. The process includes: Global Features , normal description feature set and anomaly description feature set Input to the control boundary module to calculate the mean of the normal descriptive characteristics. The formula is: , Where, Indicates the normal description template index; Construct the feature similarity loss function, the formula is: , Where, represents the natural exponential function, is the temperature coefficient, represents the negative normalized dot product similarity, Represents a feature vector in the template-type exception description feature, represents a feature vector in the evolvable anomaly template; the loss Enforce global features of normal images With normal descriptive feature mean The similarity is higher than its similarity to the abnormal description features; Implement multi-prototype space constraints and calculate the mean of anomaly description features. The formula is: , Calculate the center of the artificial anomaly template using the formula: , Calculate the center of the evolvable abnormal template using the formula: , Construct a dual distance loss function, the formula is: , in, Represents global features The expected value of represents the square of the Euclidean distance, is the distance boundary threshold, To distribute the alignment strength coefficient, the first term constrains the global features of the normal image distance The Euclidean distance ratio Nearly at least Unit, the second term narrows the distribution difference between the two types of abnormal prototypes; The overall optimization goal is , after updating the evolvable parameters through gradient descent, the optimized normal description set { }, and by exception semantic templates and An exception description set consisting of an evolvable exception template , the normal description feature mean and the mean of the anomaly descriptive features Stored in the reference library S.
6. The cross-modal text semantic-driven few-shot defect recognition method according to claim 5, characterized in that: The steps of combining the global features of the normal image with the normal description features and the abnormal description features to generate the normal feature vector and the abnormal feature vector include: Global Features Through the linear projection layer Mapped to , so that the global features are aligned with the normal description features and abnormal description feature dimensions; For each normal description feature Performing cross-modal fusion: using a gating mechanism Generate normal eigenvectors , the formula is: , Where, represents element-wise product, is the evolvable parameter matrix, Sigmoid activation; Describe the characteristics of each anomaly Calculating cross-modal attention gating weights , is an evolvable parameter; hybrid features are generated by gating weights , all After concatenation along the channel dimension, it passes through the fully connected layer Projected as the final abnormal feature vector , the formula is: ; Store normal eigenvectors and abnormal feature vector To the reference library .
7. The method for identifying defects with a small number of samples based on cross-modal text semantics according to claim 6, characterized in that: The steps of obtaining a synthetic abnormal image using the abnormal feature vector include: Using abnormal feature vectors Constructing a set of abnormal feature vectors ,Will Input into the decoding module of the pre-trained text-image generator to obtain a tensor , represents a pre-trained image-text generator that generates images based on text features. Represents the number of normal samples in the training data set, H represents the height of the synthetic abnormal image finally output by the generator, and W represents the width of the synthetic abnormal image finally output by the generator; When the network reaches the final layer, the pixel value range of the tensor z is constrained by the hyperbolic tangent function, and the original output is compressed to The interval is expressed as: , Where, It is represented as the output tensor after compression by the hyperbolic tangent function, that is, the pixel value of the synthesized image; Then, the values are mapped to the [0,255] range of the standard RGB image through linear transformation to obtain the detection samples containing the synthetic abnormal image. , the formula is: , Through the constraints Achieve resolution from Gradually improve to ,in The initial feature map base size parameter for the input generation network, For the The upsampling ratio of the layer deconvolution, is the network depth; For each sample and each exception description ] perform the generation operation independently, and the final generated set of synthetic abnormal images is expressed as: , It represents the synthetic abnormal image generated by combining the nth normal prompt template corresponding to the kth normal image sample and the mth abnormal description.
8. The cross-modal text semantic-driven few-shot defect recognition method according to claim 7, characterized in that: In the testing phase, the steps of extracting global image features and local image features of the image to be tested include: The normalized RGB query image is denoted as , where the resolution is fixed at H=W=256, Input image encoder, extract features through the multi-head self-attention mechanism of 12-layer Transformer block to obtain global image features , the formula is: , Where, represents a nonlinear activation function, is the original global semantic vector output by the image encoder, is the evolvable orthogonal projection matrix, Representation layer normalization; Will Input the multi-scale feature extractor, which consists of a 4-level convolution block, and obtains four sets of local feature maps, namely: , , , , Where, 、 、 、 They represent the first-level convolution block, the second-level convolution block, the third-level convolution block, and the fourth-level convolution block respectively. Each convolution block contains three 3×3 convolution layers and ReLU activations. For each level of local feature map Perform spatial dimension expansion and L2 normalization to generate a local image feature set. The formula is: , Where, , Indicates the Level feature map in spatial position The eigenvector at represents the Euclidean norm of a vector; In order to align with the local feature dimensions used in the training phase, the local image feature set Through the linear mapping layer Projecting to the feature space with the same dimension and semantic space as the text features extracted by CLIP, the final local image feature set is obtained = , where each local eigenvector is normalized on the unit Euclidean sphere.
9. The cross-modal text semantic-driven few-shot defect recognition method according to claim 8, characterized in that: The steps of calculating the distance between the normal feature vector and the abnormal feature vector and the global image features of the image to be detected using the nearest neighbor algorithm and generating the image-level anomaly score based on the distance difference include: Calculate global features separately and store in reference library Normal eigenvectors in , abnormal feature vector The cosine similarity between them is used to evaluate the similarity between the test image and the training features. The formulas are: , ; Where, Represents the L2 norm of the vector; Calculate the image-level anomaly score using the formula: 。 10. The cross-modal text semantic-driven few-shot defect recognition method according to claim 9, characterized in that: The steps of obtaining an image-level anomaly detection score by combining the image-level anomaly score and the prediction result and comparing the image-level anomaly score with a preset threshold to obtain a judgment result include: Absolute value weighted fusion and classification probability The fractional operation is used to obtain the image-level anomaly detection score , the formula is: , Where, represents the fusion weight coefficient, Represents the binary classification prediction result; When the image-level anomaly score When the image to be detected is determined to be an abnormal image, represents the classification decision threshold determined by receiver operating characteristic curve analysis.
11. The cross-modal text semantic-driven few-shot defect recognition method according to claim 10, characterized in that: When the image to be detected is an abnormal image, the process of detecting the predicted abnormal segmentation map of the abnormal image includes: For each spatial position in the local feature map P The eigenvector of , where the eigenvector of each spatial position is Corresponding to a local area of the image, calculate The cosine similarity with the mean of normal description features and the mean of abnormal description features is as follows: , Where, is the temperature parameter, represents the vector inner product; Represents the normalized local features represents the mean of the normal descriptive features in the reference library S, represents the mean value of the abnormal description features in the reference library S; pass Constructing a preliminary semantic anomaly score map , each element in the semantic anomaly score graph ,The larger the value, the higher the probability of regional anomaly; And select the reference library S The minimum similarity value of feature vectors at the same level , the formula is: , Where, express A local visual feature vector in Indicates the local features of the test image currently being processed The feature map level; The minimum similarity value Mapped to anomaly score, the formula is: , By measuring the minimum matching degree between the local features of the query image and the local features of the normal image, the similarity range [-1,1] is linearly converted to [0,1]. The larger the value, the higher the probability of abnormality. All abnormal scores are combined to form an abnormal heat map. , spatial resolution and input feature map Consistency directly reflects the degree to which each region deviates from the normal pattern; calculate and The nonlinear sum fusion number is used as the predicted anomaly score map , the formula is: , Where, Represents the predicted anomaly score map The value of each spatial location; The predicted anomaly score graph Perform multi-scale morphological optimization and sub-pixel boundary refinement operations, where the predicted anomaly score map first generates an initial binary mask through dynamic threshold segmentation: using the improved OTSU algorithm to calculate the adaptive segmentation threshold , the formula is: , Where, is the inter-class variance of the grayscale histogram of the heat map, is the local contrast compensation factor, is a sliding window, is the local gradient amplitude, is the gradient operator, and denote the horizontal and vertical partial derivatives respectively, is the total number of pixels; If and only if When the binary mask =1 indicates abnormality, otherwise =0 means normal, generate initial mask .
Citation Information
Patent Citations
Cable defect detection method and system based on large model
CN118115483A
Small-sample industrial anomaly detection method based on cross-modal adaptive interaction
CN119989247A
Zero-sample-driven dialogue type industrial defect detection system and method
CN120107190A
Metal defect identification method trained by few training samples
CN120374598A
Method for automatically generating concrete dam defect image description on basis of graph attention network
WO2023241272A1
Cited By
Forging defect identification data processing method and system based on surface migration prediction
CN121074028A
Geological disaster emergency decision-making method and system based on adaptive generative AI
CN121352010A
A Geological Disaster Emergency Decision-Making Method and System Based on Adaptive Generative AI
CN121352010B
Multi-source heterogeneous medical data fusion and intelligent diagnosis method
CN121545724A
Cross-domain few-sample forest land vegetation adaptive feature recognition and extraction system and method
CN121837931A