Small sample industrial defect image segmentation method and system based on text background perception enhancement, and storage medium
Through the CLIP model and joint activation graph optimization strategy, the semantic information of prospects and backgrounds is used to solve the problem of scarcity and labeling difficulties in industrial defect detection, and high-precision and stable defect segmentation are achieved, which is suitable for a variety of industrial scenarios.
Patent Information
- Application Number
- CN202511086907.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-08-05
AI Technical Summary
The existing industrial defect detection methods are difficult to achieve high-precision and stable defect segmentation when defect samples are scarce, labeling costs are high and background information is insufficient, especially in complex backgrounds, model generalization capabilities are limited.
A small sample industrial defect image segmentation method based on text background perception enhancement is adopted, multi-level visual features and text embedding vectors are extracted through the CLIP model, and a coarse-grained activation map is generated. Combined with static refinement and dynamic refinement strategies, the few sample segmentation model is guided to segment defect areas, making full use of the semantic co-occurrence relationship between foreground and background.
It significantly improves the accuracy and efficiency of defect detection, reduces the cost of data labeling, and achieves high robustness and adaptability with very few labeling samples. It is suitable for a variety of industrial visual defect detection scenarios.
Smart Images

Figure CN120580259A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and industrial automation detection, and specifically relates to a small sample industrial defect image segmentation method, system and storage medium based on text background perception enhancement. Background Art
[0002] Industrial defect detection is a critical component of intelligent manufacturing systems, with its accuracy and stability directly impacting product quality control and the level of production line automation. In recent years, deep learning methods have made significant progress in this field, with convolutional neural networks (CNNs) and Transformer architectures in particular demonstrating excellent performance in defect recognition tasks. However, industrial defect data often suffers from the following characteristics: first, defect samples are scarce, making large-scale annotated data difficult to obtain; second, defect morphology is complex and varied, exhibiting significant uncertainty and domain specificity; and third, the high cost of annotation and its strong reliance on specialized knowledge make traditional fully supervised learning methods difficult to apply on a large scale in real industrial environments.
[0003] To overcome these issues, researchers have introduced the Few-Shot Learning (FSL) mechanism, which utilizes limited labeled samples to identify new defect types. Furthermore, the recent development of multimodal pre-training models, particularly the breakthrough of the CLIP model in joint image-text modeling, has provided a new opportunity to introduce language guidance to industrial vision tasks. However, existing methods often focus on text modeling of foreground categories, overlooking the crucial role of background information in enhancing the model's discriminative ability in defect segmentation. Numerous limitations remain. First, relying on textual cues from a single foreground category fails to provide sufficient semantic discriminative power, especially in complex backgrounds with visual interference or semantic overlap, which can easily lead to model misjudgments. Second, the potential information value of background categories is overlooked, failing to leverage the semantic co-occurrence relationship between background and foreground to enhance the model's contextual understanding. Third, the quality of activation map generation is poor, often exhibiting insufficient foreground response or excessive background response, resulting in unstable segmentation results and limited generalization. Summary of the Invention
[0004] The purpose of the present invention is to address the above-mentioned problems and propose a small sample industrial defect image segmentation method, system and storage medium based on text background perception enhancement, which greatly improves the accuracy and efficiency of defect detection, realizes small sample and low-cost detection, and has strong generalization ability.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] The present invention proposes a small sample industrial defect image segmentation method based on text background perception enhancement, which includes the following steps:
[0007] S1. Construct corresponding support-query pairs based on industrial defect image sets ,in, For the query image, the support set , For the Support images, For the Mask labels for the supported images, The number of supported images;
[0008] S2. Build a text context-aware enhancement model to perform the following operations:
[0009] S21, respectively obtaining a foreground text and multiple background texts of a defective area of each supporting image;
[0010] S22, input the query image into the image encoder of the CLIP model, extract the multi-level visual features and the multi-level attention weight matrix, and input the foreground text and each background text of the support image of the corresponding category into the text encoder of the CLIP model, and obtain the foreground text embedding vector and each background text embedding vector;
[0011] S23, generating a corresponding coarse-grained activation map using the Grad-CAM method based on the multi-level visual features of the query image, the foreground text embedding vector, and the background text embedding vectors;
[0012] S24, obtaining a static refined activation map and a dynamic refined activation map according to the coarse-grained activation map, the multi-level visual features, and the multi-level attention weight matrix;
[0013] S25. Input the query image, the corresponding support image, and the mask label into the few-shot segmentation model, and use the static refined activation map and the dynamic refined activation map as attention guidance signals;
[0014] S3. Use all support-query pairs to train a text context-aware enhancement model;
[0015] S4. Input the query image to be segmented, the corresponding support image and the mask label into the trained text background perception enhancement model, and obtain the output result of the few-shot segmentation model as the segmentation prediction result of the corresponding query image.
[0016] Preferably, the query image is an industrial defect image in the industrial defect image set, the support image is an industrial defect image with pre-labeled defect categories, and the mask label of the support image is a binary image or a multi-class labeled image corresponding to all pixels in the support image.
[0017] Preferably, the foreground text and multiple background texts of the defect area of each supporting image are generated using the DeepSeek-R1 large language model according to the preset question-answer template; multi-level visual features , For the Layer visual features, multi-level attention weight matrix , For the layer attention weight matrix, , is the total number of layers of the image encoder, To query the height of the image, The width of the query image.
[0018] Preferably, the Grad-CAM method is used to generate the corresponding coarse-grained activation map, as follows:
[0019] S231, calculate the multi-level visual features The cosine similarity between the visual features of each spatial position of the layer visual feature and the foreground text embedding vector and the background text embedding vector, is the total number of layers of the image encoder;
[0020] S232, normalizing each cosine similarity using a softmax function to form a corresponding probability vector;
[0021] S233, the cross entropy loss between the probability vector and the true label is used as the coarse-grained activation map to generate the loss function, where the foreground of the true label is 1 and the background is 0;
[0022] S234, calculate the coarse-grained activation map to generate the loss function for the The first layer of visual features Spatial position in the channel The partial derivative of the visual feature of the corresponding channel is the spatial position The gradient of the visual features, is the height direction index of the query image, is the width index of the query image, , , To query the height of the image, is the width of the query image;
[0023] S235, the All gradients of each channel of the visual features of the layer are globally averaged and pooled to obtain the global gradient of the corresponding channel, and the global gradient of each channel is used to average the gradient of the first layer. The visual features of the corresponding channels of the layer visual features are weighted summed, and then the coarse-grained activation map is generated through the ReLU activation function.
[0024] Preferably, a static refined activation map and a dynamic refined activation map are obtained according to the coarse-grained activation map, the multi-level visual features, and the multi-level attention weight matrix, as follows:
[0025] S241. Obtaining a static refinement matrix and a dynamic refinement matrix based on a joint refinement strategy according to the multi-level visual features and the multi-level attention weight matrix;
[0026] S242. Perform matrix multiplication operations on the coarse-grained activation map with the static refinement matrix and the dynamic refinement matrix, respectively, to obtain a static refinement activation map and a dynamic refinement activation map, respectively.
[0027] Preferably, after obtaining the static refined activation map and the dynamic refined activation map based on the coarse-grained activation map, the multi-level visual features, and the multi-level attention weight matrix, a pseudo label is generated from the dynamic refined activation map according to a preset threshold, that is, when the activation value of each spatial position in the dynamic refined activation map is greater than or equal to the preset threshold, it is considered to be a foreground area of the pseudo label; otherwise, it is considered to be a background area of the pseudo label;
[0028] The joint refinement strategy includes static refinement strategy and dynamic refinement strategy, where:
[0029] The static refinement strategy performs the following operations:
[0030] The attention weight matrix of each layer in the multi-level attention weight matrix is averaged and the normalized matrix is obtained using the Sinkhorn algorithm;
[0031] Perform matrix multiplication operation on the normalized matrix and the transpose of the normalized matrix to obtain a high-order optimized matrix;
[0032] Take the larger value of the high-order optimization matrix and the normalized matrix at each spatial position to generate a static refinement matrix;
[0033] The dynamic refinement strategy performs the following operations:
[0034] Concatenate all the visual features of the multi-level visual features in the channel dimension and perform convolution operation Obtaining intermediate layer features;
[0035] Perform a shape reshaping operation on the intermediate layer features to form a reshaped feature, and then perform a matrix multiplication operation on the reshaped feature and the transpose of the reshaped feature to form an initial dynamic refinement matrix;
[0036] Calculate the difference between the attention weight matrix of each layer in the multi-level attention weight matrix and the corresponding element in the initial dynamic refinement matrix, and add the differences of all elements in the corresponding layer to form the fusion feature of the corresponding layer;
[0037] The fusion features of each layer are averaged to obtain the average matrix, and the attention weight matrix with higher than the average matrix in the multi-level attention weight matrix is selected as the middle layer attention matrix;
[0038] The selected middle layer attention matrix is averaged over the channels to obtain the middle layer optimization matrix;
[0039] The initial dynamic refinement matrix is multiplied pixel by pixel with the intermediate layer optimization matrix to obtain the dynamic refinement matrix of the current training round, and the cross entropy loss between the dynamic refinement matrix of the current training round and the pseudo label is used as the loss function of the dynamic refinement matrix of the next training round.
[0040] Preferably, the segmentation loss of the few-shot segmentation model is the cross entropy loss between the segmentation prediction result of the query image and the true mask label.
[0041] Preferably, the few-shot segmentation model is a PFENet model or a HDMNet model.
[0042] A small sample industrial defect image segmentation system based on text background perception enhancement includes a data acquisition module, a model building module, a model training module and a prediction module, wherein:
[0043] Data acquisition module, used to construct corresponding support-query pairs based on industrial defect image sets ,in, For the query image, the support set , For the Support images, For the Mask labels for the supported images, The number of supported images;
[0044] The model building module is used to build a text context-aware enhancement model to perform the following operations:
[0045] respectively obtaining a foreground text and a plurality of background texts of a defect area of each supporting image;
[0046] The query image is input into the image encoder of the CLIP model, and multi-level visual features and multi-level attention weight matrices are extracted accordingly. The foreground text and background text of the support image of the corresponding category are input into the text encoder of the CLIP model, and the foreground text embedding vector and background text embedding vector are obtained accordingly.
[0047] According to the multi-level visual features of the query image, the foreground text embedding vector and the background text embedding vector, the Grad-CAM method is used to generate the corresponding coarse-grained activation map;
[0048] Obtain static refined activation maps and dynamic refined activation maps based on the coarse-grained activation map, multi-level visual features, and multi-level attention weight matrix;
[0049] The query image, the corresponding support image, and the mask label are fed into the few-shot segmentation model, and the static refined activation map and the dynamic refined activation map are used as attention guidance signals.
[0050] A model training module is used to train a text context-aware enhancement model using all support-query pairs;
[0051] The prediction module is used to input the query image to be segmented, the corresponding support image and the mask label into the trained text background perception enhancement model, and obtain the output result of the few-shot segmentation model as the segmentation prediction result of the corresponding query image.
[0052] A storage medium for small sample industrial defect image segmentation based on text background perception enhancement is used to store a computer program. When the computer program is executed by a processor, it implements any of the above small sample industrial defect image segmentation methods based on text background perception enhancement.
[0053] Compared with the prior art, the present invention has the following beneficial effects:
[0054] This paper addresses the problems of sample scarcity, difficulty and high cost in labeling, and insufficient model generalization in industrial defect detection tasks, especially in the context of traditional few-sample learning methods ignoring background semantic information, resulting in poor robustness of segmentation results. This paper proposes a small-sample industrial defect image segmentation method based on text background perception enhancement. By integrating a multimodal language-visual modeling mechanism with a joint activation map optimization strategy, the generalization and semantic discrimination capabilities of the industrial defect segmentation model under the condition of scarce labeled samples are significantly improved, greatly improving the accuracy and efficiency of defect detection. Specifically:
[0055] 1) Based on a collection of industrial defect images, corresponding support-query pairs are constructed. A small number of annotated support images and corresponding mask labels are read, and a query image of unknown category is input. A large language model (such as DeepSeek-R1) is combined with preset question-answering templates to automatically generate foreground and background text related to the defect. Based on the image modality, a dual foreground-background semantic guidance mechanism is introduced to enrich the model's semantic perception information, guiding the model to achieve more detailed perception of the target defect area, thereby enhancing the model's understanding of semantic structure.
[0056] 2) The query image is input into the CLIP model's image encoder to extract multi-scale image features (multi-level visual features and multi-level attention weight matrices). The foreground and background text are simultaneously input into the CLIP model's text encoder to extract the corresponding foreground text embedding vectors and background text embedding vectors. The Grad-CAM method is then used to generate an initial activation map (coarse-grained activation map). This involves jointly extracting visual and language features through the CLIP multimodal model. The activation regions are further refined through a joint optimization strategy to avoid insufficient foreground response or background interference. Under the dual foreground-background semantic guidance mechanism, multimodal image-text alignment modeling and joint activation map optimization are combined to enhance foreground response and background noise suppression capabilities based on static noise suppression constraints and dynamic pseudo-mask supervision. This improves the model's perception of the target area and its ability to suppress background noise, even with only a small number of labeled samples. This in turn enhances the model's segmentation capability, achieving high-precision recognition of industrial defect targets and effective filtering of background noise, significantly improving the robustness and adaptability of the industrial defect segmentation task.
[0057] 3) By fully exploiting the semantic co-occurrence relationship between foreground and background, this method overcomes the problem of existing few-shot segmentation methods that rely solely on foreground cues and ignore background information. It can achieve high-precision segmentation of industrial defect areas with only a very small number of labeled samples, significantly reducing data annotation costs. It is applicable to a variety of industrial visual defect detection scenarios and has good engineering deployment and generalization capabilities.
[0058] 4) The static and dynamic refined activation maps are used as attention guidance signals and fed into the few-shot segmentation model along with the query image, corresponding support image, and mask label for training. This training guides the model to segment the defective areas of the query image until the model converges. The model weights generated by the training can then be used to segment industrial defect images.
[0059] The present invention does not require a large amount of labeled data and can achieve highly robust and high-precision industrial defect segmentation by relying on only a few supporting images. Through the deep fusion of text and image information, it fully utilizes the prior information provided by background semantics to guide the model to perform structural perception and background suppression in defect recognition, significantly improving the accuracy, stability and industrial applicability of small sample segmentation. It has a simple structure and strong adaptability, and is particularly suitable for small sample detection tasks in a variety of complex industrial scenarios, and has good practical engineering application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 This is a flow chart of the small sample industrial defect image segmentation method based on text background perception enhancement of the present invention;
[0061] Figure 2Schematic diagram of the structure of the text background perception enhancement model of the present invention;
[0062] Figure 3 is a flow chart of the joint refinement strategy of the present invention;
[0063] Figure 4 Comparison diagrams of support images, mask labels of support images, query images, manual annotation results of query images, and segmentation prediction results of query images using the method of the present invention, wherein (A) is the support image, (B) is the mask label of the support image, (C) is the query image, (D) is the manual annotation result of the query image, and (E) is the segmentation prediction result of the query image using the method of the present invention. DETAILED DESCRIPTION
[0064] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0065] It should be noted that when a component is referred to as being "connected" to another component, it may be directly connected to the other component or there may be an intermediate component. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art in the art of this application. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0066] Example 1:
[0067] like Figures 1-4 As shown in FIG, a small sample industrial defect image segmentation method based on text background perception enhancement includes the following steps:
[0068] S1. Construct corresponding support-query pairs based on industrial defect image sets ,in, For the query image, the support set , For the Support images, For the Mask labels for the supported images, The number of supported images.
[0069] In one embodiment, the query image is an industrial defect image in an industrial defect image set, the support image is an industrial defect image with defect categories pre-labeled, and the mask label of the support image is a binary image or a multi-class labeled image corresponding to all pixels in the support image.
[0070] Among them, the support-query pair constitutes the basic task unit in the small sample segmentation task, and the organization form is :Each query image corresponds to a set of support images containing defect category annotations, that is, a support-query pair consisting of one or more support images and a single query image. In terms of quantity, through Supporting images ( -shot) together provide semantic guidance for the same query image, thereby achieving accurate segmentation of defect areas of the same category in the query image.
[0071] In small-sample segmentation tasks, support images are image samples that are pre-labeled with defect categories. They typically contain only a small number of representative defect instances and provide semantic guidance by providing limited category information, guiding the subsequently established text context-aware enhancement model to identify defects of the same category in the query image. Mask labels are binary or multi-class labeled images that indicate whether each pixel in the support image belongs to the defect region. Specifically, the representation of different categories is unified by the value of each pixel in the support image, that is, the same category is represented by the same pixel color, which is used to provide accurate spatial location and category information. The query image is the target image that the text context-aware enhancement model needs to reason about and segment based on the category clues provided by the support image. Its defect category is consistent with the support image but without any explicit labeling information. The query image also reflects the generalization reasoning ability of the text context-aware enhancement model (segmentation model) under weak supervision and is key to evaluating the ability of the text context-aware enhancement model to recognize new categories under low data volume conditions.
[0072] Furthermore, defect category labeling of support images involves pixel-level segmentation and annotation of defect regions within the support images. Specifically, each pixel in the defect region within the industrial defect image is labeled as foreground or background. This is typically represented as a two-dimensional matrix of the same size as the industrial defect image, where each element is a 1 (or a positive integer category number) indicating that the pixel belongs to the defect region (foreground), and a 0 indicating that the pixel is background. This industrial defect image set can specifically represent data from multiple industrial scenarios, such as metal processing, semiconductor packaging, photovoltaic glass, textile materials, and automotive manufacturing. By collecting image data containing defect features (such as steel plate cracks, wafer defects, glass breakage, and fabric noise), it provides the necessary data foundation for the subsequent construction of support-query pairs for defect segmentation. This can be applied to related fields or generalized to other fields.
[0073] S2. Build a text context-aware enhancement model to perform the following operations:
[0074] S21, respectively obtaining a foreground text and multiple background texts of a defective area of each supporting image;
[0075] S22, input the query image into the image encoder of the CLIP model, extract the multi-level visual features and the multi-level attention weight matrix, and input the foreground text and each background text of the support image of the corresponding category into the text encoder of the CLIP model, and obtain the foreground text embedding vector and each background text embedding vector;
[0076] S23, generating a corresponding coarse-grained activation map using the Grad-CAM method based on the multi-level visual features of the query image, the foreground text embedding vector, and the background text embedding vectors;
[0077] S24, obtaining a static refined activation map and a dynamic refined activation map according to the coarse-grained activation map, the multi-level visual features, and the multi-level attention weight matrix;
[0078] S25. Input the query image, the corresponding support image and the mask label into the few-shot segmentation model, and use the static refinement activation map and the dynamic refinement activation map as attention guidance signals.
[0079] In one embodiment, the foreground text and multiple background texts of the defect area of each supporting image are generated using the DeepSeek-R1 large language model according to the preset question-answer template; multi-level visual features , For the Layer visual features, multi-level attention weight matrix , For the layer attention weight matrix, , is the total number of layers of the image encoder, To query the height of the image, The width of the query image.
[0080] For each supporting image, the DeepSeek-R1 large language model is used based on the preset question-answering template to automatically generate foreground text and Background text (text prompt words and text feature generation). The foreground text describes the typical semantic features of the target defect area, and the background text automatically infers the background information with semantic co-occurrence relationship through the preset question-answer template, providing multimodal semantic support for subsequent segmentation tasks. Figure 2 As shown, the preset question and answer template is:
[0081] Which objects are most likely to appear in an image containing { },but are not part of { }?
[0082] Which objects are most likely to be the background of { }?
[0083] Combine the answers to the two questions and output in['###'] format.
[0084] After inputting the preset question-answer template into the DeepSeek-R1 large language model, a list of background categories associated with the foreground category will be returned, that is, the two question answers will be merged and the corresponding English words will be output, such as ['Intact Steel', 'SmoothMetal Surface', 'Brushed Metal', 'Polished Stainless Steel', 'Coated Steel','Flat Steel Sheet', 'Clean Metal Plate', 'Uniform Steel Surface', 'NormalRolled Surface', ' Well-Finished Steel'], which are used to describe the background objects that appear most frequently in the image containing the foreground target.
[0085] Specifically, the output foreground text is A clear origami { }, { The query image's foreground text is replaced by the text of the defect type or target object of interest. For example, if the foreground text is: A clear origami {Steel Laceration Defect}, the output background text is A clear origami {###}. [###] represents the list of background categories associated with the foreground category, generated by the DeepSeek-R1 large language model. The background text is: A clear origami {Intact Steel}, A clear origami {Clean Metal Plate}, …, A clear origami {Normal Rolled Surface}. This preset question-answering template can be used to obtain background text that co-occurs with the query image's foreground text with high probability.
[0086] The CLIP model consists of an image encoder and a text encoder, which are used to map images and text into a unified semantic space. This technology is well known to those skilled in the art and will not be described in detail here. For example, the CLIP model in this embodiment is based on the following literature: Radford A, Kim JW, Hallacy C, et al. Learning transferable visual models from natural language supervision [C] / / International conference on machine learning. PmLR, 2021: 8748-8763.
[0087] The query image is input into the image encoder of the CLIP model, corresponding to the extraction of multi-level visual features. and multi-level attention weight matrix At the same time, the foreground text and background text of the support image of the corresponding category are input into the text encoder of the CLIP model, and the corresponding foreground text embedding vector and background text embedding vector are obtained. The formula is as follows:
[0088]
[0089] in, is the image encoder of the CLIP model, is the text encoder of the CLIP model, To support foreground text on images, To support the image Background text, is the foreground text embedding vector, For the background text embedding vectors, , The amount of background text.
[0090] In one embodiment, the Grad-CAM method is used to generate the corresponding coarse-grained activation map, as follows:
[0091] S231, calculate the multi-level visual features The cosine similarity between the visual features of each spatial position of the layer visual feature and the foreground text embedding vector and the background text embedding vector, is the total number of layers of the image encoder;
[0092] S232, normalizing each cosine similarity using a softmax function to form a corresponding probability vector;
[0093] S233, the cross entropy loss between the probability vector and the true label is used as the coarse-grained activation map to generate the loss function, where the foreground of the true label is 1 and the background is 0;
[0094] S234, calculate the coarse-grained activation map to generate the loss function for the The first layer of visual features Spatial position in the channel The partial derivative of the visual feature of the corresponding channel is the spatial position The gradient of the visual features, is the height direction index of the query image, is the width index of the query image, , , To query the height of the image, is the width of the query image;
[0095] S235, the All gradients of each channel of the visual features of the layer are globally averaged and pooled to obtain the global gradient of the corresponding channel, and the global gradient of each channel is used to average the gradient of the first layer. The visual features of the corresponding channels of the layer visual features are weighted summed, and then the coarse-grained activation map is generated through the ReLU activation function.
[0096] Specifically, if Figure 2 As shown, the coarse-grained activation map The generation process is as follows:
[0097] 1) Calculate the first The cosine similarity between the visual features of each spatial position of the layer visual feature and the foreground text embedding vector and the background text embedding vector is as follows:
[0098]
[0099] in, is the text embedding vector, that is, the foreground text embedding vector or the background text embedding vector, For the Layer visual features spatial location Visual features With text embedding vector cosine similarity (response strength), is the Euclidean norm;
[0100] 2) Normalize each cosine similarity through the softmax function to form a corresponding probability vector ;
[0101] 3) The probability vector and the true label The cross entropy loss is used as a coarse-grained activation map Loss function during generation , the formula is as follows:
[0102]
[0103] in, is the cross entropy loss, the true label It consists of a one-hot encoding of the category to which the query image belongs, such as 1 for foreground and 0 for background;
[0104] 4) For the Grad-CAM method, the calculation of its weight depends on the target category score in the multi-level visual features. The gradient of the visual features of the first layer (that is, the last layer, the image encoder of the CLIP model has 12 layers in total), then The first layer of visual features Spatial position in the channel The gradient of the visual feature , the formula is as follows:
[0105]
[0106] in, Loss function representing coarse-grained activation maps For the first Layer visual features No. Spatial position in the channel The partial derivative of the visual feature is The first layer of visual features Spatial position in the channel The gradient of the visual feature , which means the coarse-grained activation map Loss function during generation For the first The first layer of visual features Spatial position in the channel The response degree of the visual feature, k=1~K, K is the The total number of channels of the layer's visual features. This gradient serves as the basis for the Grad-CAM weights, guiding the model to focus on areas that contribute significantly to specific categories.
[0107] 5) Generate coarse-grained activation map , the formula is as follows:
[0108]
[0109] in, For the The global gradient of the channels is expressed as , For the The first layer of visual features The visual features of the channels are represented by ReLU, which is a ReLU activation function used to retain areas with positive contributions (positive values).
[0110] In one embodiment, a static refined activation map and a dynamic refined activation map are obtained based on the coarse-grained activation map, the multi-level visual features, and the multi-level attention weight matrix, as follows:
[0111] S241. Obtaining a static refinement matrix and a dynamic refinement matrix based on a joint refinement strategy according to the multi-level visual features and the multi-level attention weight matrix;
[0112] S242. Perform matrix multiplication operations on the coarse-grained activation map with the static refinement matrix and the dynamic refinement matrix, respectively, to obtain a static refinement activation map and a dynamic refinement activation map, respectively.
[0113] In order to improve the accuracy and segmentation guidance ability of the coarse-grained activation map, a joint refinement strategy based on the static refinement strategy and the dynamic refinement strategy is designed. The static refinement strategy uses the background suppression constraint to obtain the static refinement matrix , suppressing background noise. The dynamic refinement strategy introduces a learnable layer to supervise the generation of the dynamic refinement matrix , improve the foreground area response. Finally, through the output matrix of the two refinement strategies, different refined high-quality activation maps are generated (static refinement activation map and is the dynamic refinement activation map ); and the current dynamic refinement activation map By threshold screening (preset threshold ) to generate the corresponding pseudo labels To further supervise the next training round to dynamically refine the matrix The generation of . Figure 2 As shown, is the static refined activation map, To dynamically refine the activation map, is the static refinement matrix, is the dynamic refinement matrix.
[0114] In one embodiment, after obtaining a static refined activation map and a dynamic refined activation map based on the coarse-grained activation map, multi-level visual features, and a multi-level attention weight matrix, a pseudo label is generated from the dynamic refined activation map according to a preset threshold. That is, when the activation value of each spatial position in the dynamic refined activation map is greater than or equal to the preset threshold, it is considered to be a foreground area of the pseudo label; otherwise, it is considered to be a background area of the pseudo label.
[0115] The joint refinement strategy includes static refinement strategy and dynamic refinement strategy, where:
[0116] The static refinement strategy performs the following operations:
[0117] The attention weight matrix of each layer in the multi-level attention weight matrix is averaged and the normalized matrix is obtained using the Sinkhorn algorithm;
[0118] Perform matrix multiplication operation on the normalized matrix and the transpose of the normalized matrix to obtain a high-order optimized matrix;
[0119] Take the larger value of the high-order optimization matrix and the normalized matrix at each spatial position to generate a static refinement matrix;
[0120] The dynamic refinement strategy performs the following operations:
[0121] Concatenate all the visual features of the multi-level visual features in the channel dimension and perform convolution operation Obtaining intermediate layer features;
[0122] Reshape the intermediate layer features to form reshaped features, and then perform matrix multiplication on the reshaped features and the transpose of the reshaped features to form an initial dynamic refinement matrix;
[0123] Calculate the difference between the attention weight matrix of each layer in the multi-level attention weight matrix and the corresponding element in the initial dynamic refinement matrix, and add the differences of all elements in the corresponding layer to form the fusion feature of the corresponding layer;
[0124] The fusion features of each layer are averaged to obtain the average matrix, and the attention weight matrix with higher than the average matrix in the multi-level attention weight matrix is selected as the middle layer attention matrix;
[0125] The selected middle layer attention matrix is averaged over the channels to obtain the middle layer optimization matrix;
[0126] The initial dynamic refinement matrix is multiplied pixel by pixel with the intermediate layer optimization matrix to obtain the dynamic refinement matrix of the current training round, and the cross entropy loss between the dynamic refinement matrix of the current training round and the pseudo label is used as the loss function of the dynamic refinement matrix of the next training round.
[0127] Specifically, if Figure 3 As shown in , the joint refinement strategy includes a static refinement strategy (static refinement branch) and a dynamic refinement strategy (dynamic refinement branch).
[0128] In the static refinement strategy, the attention weight matrix of each layer in the multi-level attention weight matrix is averaged (the average weight matrix is obtained ) Use the Sinkhorn algorithm (Sinkhorn) to obtain the normalized matrix According to the normalized matrix Get high-order optimization matrix , the formula is as follows:
[0129]
[0130] in, is the normalized matrix The transpose of is matrix multiplication.
[0131] Then the high-order optimization matrix and the normalized matrix In the corresponding value of each spatial position, the larger one is selected as the output, which can fuse the information contained in the two matrices to generate the final static refinement matrix. .
[0132] Dynamic Refinement Matrix To be a learnable dynamic matrix, first, we need to obtain the intermediate layer features based on all the layer visual features in the multi-level visual features (the 12-layer visual features of the image encoder of the CLIP model). , the formula is as follows:
[0133]
[0134] in, is the convolution operation, To stitch in the channel dimension;
[0135] According to the characteristics of the middle layer Calculate the initial dynamic refinement matrix , the formula is as follows:
[0136]
[0137] in, Represents the shape reshaping operation, using the permute function, is transposed;
[0138] Next, the initial dynamic refinement matrix is refined by the attention weight matrix of each layer Optimize. First, calculate the Layer attention weight matrix With the initial dynamic refinement matrix The difference of the corresponding elements in the corresponding layer is added to form the fusion feature of the corresponding layer , to indicate and The actual similarity is then calculated by channel to get the average matrix , only select the actual similarity higher than the average matrix The attention weight matrix is used as the intermediate layer attention matrix (Contrast and retain the high value index.) Finally, the selected middle layer attention matrix is averaged over the channels to obtain the middle layer optimization matrix .Will and Pixel-by-pixel multiplication to obtain the final dynamic refinement matrix . Loss function of dynamic refinement matrix , the formula is as follows:
[0139]
[0140] in, is the cross entropy loss, is matrix multiplication.
[0141] In the joint refinement strategy, the static refinement strategy uses static statistics to suppress interference from non-target areas, and the dynamic refinement strategy uses the activation map generated by the previous optimization step as a pseudo-mask to enhance the response of the foreground area. After joint optimization, the two output more refined and robust static refinement activation maps and dynamic refinement activation maps, guiding the segmentation model to focus on the target defect area more accurately, with strong stability, versatility and engineering application prospects.
[0142] In one embodiment, the segmentation loss of the few-shot segmentation model is the cross entropy loss between the segmentation prediction result of the query image and the true mask label.
[0143] Specifically, the segmentation loss of the few-shot segmentation model is , the formula is as follows:
[0144]
[0145] in, is the segmentation prediction result of the query image, is the true mask label.
[0146] In one embodiment, the few-shot segmentation model is a PFENet model or a HDMNet model.
[0147] Among them, the few-shot segmentation model Refers to any segmentation model in the field of few-shot segmentation, such as the PFENet model, the HDMNet model, etc., which is used to extract features from the support image and fuse them from the static refined activation map and dynamically refined activation maps Structural semantic information of the ,guides the few-shot segmentation model Segment the defect area in the query image. Static Refinement Activation Map and dynamically refined activation maps As a few-shot segmentation model The attention guidance signal is used to enhance the response of the defect-related areas in the query image and suppress background interference information, that is, to enhance the focusing ability of the few-shot segmentation model on the target defect area, guide it to complete the segmentation prediction of the target defect area in the query image, and output the segmentation prediction result of the corresponding query image.
[0148] S3. Use all support-query pairs to train a text context-aware enhancement model.
[0149] S4. Input the query image to be segmented, the corresponding support image and the mask label into the trained text background perception enhancement model, and obtain the output result of the few-shot segmentation model as the segmentation prediction result of the corresponding query image.
[0150] Among them, the segmentation prediction results of the query image can accurately mark the shape, location and boundary of the defect, and are used in industrial inspection systems to highlight abnormal areas in real time, facilitating the subsequent calculation of the defect's morphological information, such as perimeter and area, while also facilitating confirmation and judgment by operators, thereby improving the interpretability of inspections.
[0151] This paper addresses the challenges of sample scarcity, labeling difficulties, and insufficient model generalization in industrial defect detection tasks. This is particularly true given the reality that traditional few-shot learning methods ignore background semantic information, resulting in poorly robust segmentation results. This paper proposes a small-sample industrial defect image segmentation method based on text-based context perception enhancement. This method automatically generates foreground and background text using the DeepSeek-R1 large language model. It also introduces a foreground-background semantic linkage mechanism based on image modality, jointly extracts visual and language features using the CLIP multimodal model, and generates an initial activation map (coarse-grained activation map) using the Grad-CAM method. The activation regions are further refined through a joint optimization strategy, guiding the segmentation model to more accurately focus on the target defect area. This method demonstrates strong stability, versatility, and promising engineering applications.
[0152] For ease of understanding, the following is detailed description through specific embodiments.
[0153] like Figure 1 As shown, the core of the method of this embodiment is to introduce a foreground-background dual semantic guidance mechanism and improve the defect segmentation accuracy under the condition of very few labeled samples through a joint activation map refinement strategy. The main steps are as follows:
[0154] Step 1: Data Reading and Support-Query Task Construction: Support images and query images are read from the industrial defect image collection to construct support-query pairs. The support images contain defect label information and can be used to provide foreground semantic guidance. The query image, however, does not contain label information and serves as the target for the model to predict. Defect category names are extracted from the support images, and foreground text prompts, such as "Steel Laceration Defect," are constructed to guide the model's focus on potential defect areas in the image.
[0155] Step 2: Generate foreground and background text: Generate background text prompts using the DeepSeek-R1 large language model. This process uses a preset question-and-answer template as input, combined with defect categories. In the specific embodiment, it uses: "Which objects are most likely to appear in an image containing {Defect}, but are not part of {Steel Laceration Defect}? and Which objects are most likely to be the background of {Steel Laceration Defect}?" to infer background keywords that co-occur with the foreground semantics but are mutually exclusive in category. Finally, the foreground and background texts are input into the CLIP model for semantic encoding and cross-modal alignment. Here, Defect represents each specific defect category in the industrial defect image set.
[0156] Step 3: The image data obtained in Steps 1 and 2, along with the foreground and background text, are fed into the CLIP model. The image encoder extracts image features, while the text encoder generates semantic embeddings. Grad-CAM is then used to generate preliminary foreground and background activation maps (coarse-grained activation maps). These maps roughly reflect possible image defects or background interference. However, due to the inaccurate response of coarse-grained activation maps, further refinement is needed to enhance their guidance and reliability.
[0157] Step 4: To optimize the coarse-grained activation map, the coarse-grained activation map is processed by the static refinement branch and the dynamic refinement branch respectively through the joint refinement strategy: In the static refinement branch, the attention weight matrix of each layer obtained by the CLIP model is first averaged through the channels and then normalized by Sinkhorn. Then, the high-order optimization matrix and the normalized matrix information are fused to generate the static refinement matrix. , to suppress abnormally high response areas and remove false activation areas; in the dynamic refinement branch, feature fusion is performed through the learnable layer (the cross entropy loss of the dynamic refinement matrix of the current training round and the pseudo label is used as the loss function of the dynamic refinement matrix of the next training round), that is, the visual features obtained by the CLIP model are optimized and learned to obtain the initial dynamic refinement matrix , and through the initial dynamic refinement matrix Guided Dynamic Refinement Matrix The generation of drives the attention update of the middle layer, thereby enhancing the response of the foreground area and optimizing the edge details. Finally, the refinement results of the two branches are fused with the coarse-grained activation map. Multiply and output high-quality refined guided activation map (static refined activation map and dynamically refined activation maps ), as the guiding input for subsequent model training. Dynamically refine the activation map At the same time, the preset threshold Processing, preserving dynamic refinement activation map Greater than or equal to the preset threshold The value of , in order to generate the corresponding supervised pseudo label To guide the next round of dynamic refinement matrix Generation.
[0158] Step 5: Apply the refined guided activation map as an attention guide signal to the feature channel of the query image to improve the model's ability to focus on defect areas. At the same time, the semantic features that support image extraction are fused with the enhanced image features in the query image, and the few-shot segmentation model is used to extract the semantic features. The final segmentation prediction result is output and the cross entropy loss is calculated between it and the true mask label corresponding to the query image to guide model learning.
[0159] Step 6: Model training. In this embodiment, the number of training rounds is set to 300. During this period, the model will automatically cycle through steps 3 to 5 and pass the set threshold parameters. To determine whether the model has converged, when the model is If no accuracy improvement is seen in the rounds of training, the training is stopped. Set to 100, otherwise the model will stop training after 300 rounds.
[0160] The specific embodiment of the present invention uses the public industrial surface defect few-sample dataset FSSD-12 as the industrial defect image set for verification. This dataset has 12 defect types, including Steel_Ws (wrinkles), Steel_Se (edge scratches), Steel_Sc (center scratches), Steel_Rp (roller marks), Steel_Ri (rust inclusions), Steel_Pk (pot knocks), Steel_Pa (repair areas), Steel_Os (oxidation spots), Steel_Op (oil stains), Steel_Ld (crack defects), Steel_Ia (inclusion areas), and Steel_Am (wear marks), with a total of more than 2,000 images. In this embodiment, the PFENet model is used as the few-sample segmentation model. At the same time, the Steel_Ws (wrinkles), Steel_Se (edge scratches), Steel_Sc (center scratches), Steel_Rp (roller marks), Steel_Ri (rust inclusions), Steel_Pk (pot knocks), Steel_Pa (repaired areas), Steel_Os (oxidation spots), and Steel_Op (oil stains) categories are used as training sets, and the Steel_Ld (crack defects), Steel_Ia (inclusion areas), and Steel_Am (wear marks) categories are used as test sets. During the training and testing process, one or five images in a certain category are randomly selected as support images, and other images in the category are used as query images. Depending on the number of support images used, the corresponding methods are 1-shot (single-sample learning) and 5-shot (five-sample learning). In addition, in the experimental design, Fold-x (x=0,1,2) was used for testing. Fold-x represents a cross-category split strategy, which is often used to simulate the generalization ability under different task distributions in small-sample segmentation. Each Fold-x represents a combination of the 12 categories of defects in the dataset into training classes (support classes) and test classes (query classes), which are used to construct the few-sample segmentation task. The segmentation accuracy index uses mIoU (mean intersection over union) as the main segmentation index to comprehensively reflect the model's recognition and segmentation capabilities in the target defect area. The 1-Shot segmentation results are shown below. Figure 4 As shown in Table 1. Figure 4 It can be seen that the method of the present invention has a very high segmentation accuracy compared with the manual labeling results, which effectively saves manpower. Figure 4 (A) shows the support image, (B) shows the mask label of the support image, (C) shows the query image, (D) shows the manual annotation result of the query image, and (E) shows the segmentation prediction result of the query image using our method. Furthermore, Table 1 shows that our method achieves optimal performance in both 1-shot (single-sample learning) and 5-shot (five-sample learning) settings, with mIoU of 66.4% and 68.7%, respectively, significantly outperforming other compared methods. For example, in the 1-shot setting, it achieves an 11.1% improvement over the traditional PFENet model. This result demonstrates the robustness and accurate segmentation capabilities of our method under sample-scarce conditions.
[0161] Table 1 Comparison of mIoU indicator detection results of different model methods on the FSSD-12 dataset
[0162]
[0163] Compared with existing segmentation model methods (such as BAM model, HSNet model, PFENet model, DCP model, and CPANet model), the method of the present invention has higher accuracy, meets the needs of actual industrial detection, and greatly saves manual labeling costs.
[0164] Example 2:
[0165] A small sample industrial defect image segmentation system based on text background perception enhancement includes a data acquisition module, a model building module, a model training module and a prediction module, wherein:
[0166] Data acquisition module, used to construct corresponding support-query pairs based on industrial defect image sets ,in, For the query image, the support set , For the Support images, For the Mask labels for the supported images, The number of supported images;
[0167] The model building module is used to build a text context-aware enhancement model to perform the following operations:
[0168] respectively obtaining a foreground text and a plurality of background texts of a defect area of each supporting image;
[0169] The query image is input into the image encoder of the CLIP model, and multi-level visual features and multi-level attention weight matrices are extracted accordingly. The foreground text and background text of the support image of the corresponding category are input into the text encoder of the CLIP model, and the foreground text embedding vector and background text embedding vector are obtained accordingly.
[0170] According to the multi-level visual features of the query image, the foreground text embedding vector and the background text embedding vector, the Grad-CAM method is used to generate the corresponding coarse-grained activation map;
[0171] Obtain static refined activation maps and dynamic refined activation maps based on the coarse-grained activation map, multi-level visual features, and multi-level attention weight matrix;
[0172] The query image, the corresponding support image, and the mask label are fed into the few-shot segmentation model, and the static refined activation map and the dynamic refined activation map are used as attention guidance signals.
[0173] A model training module is used to train a text context-aware enhancement model using all support-query pairs;
[0174] The prediction module is used to input the query image to be segmented, the corresponding support image and the mask label into the trained text background perception enhancement model, and obtain the output result of the few-shot segmentation model as the segmentation prediction result of the corresponding query image.
[0175] Specifically, based on the same inventive concept as in Example 1, the data acquisition module in this embodiment is used to load support images and query images from local or remote storage to construct corresponding support-query pairs. Model building module: constructs foreground prompt words based on the defect categories marked in the support images, and at the same time generates background prompt words that have a co-occurrence relationship with the foreground semantics through the DeepSeek-R1 large language model combined with a preset question-answer template. In addition, before generating foreground text and background text, the model building module also supports image preprocessing operations such as unified resizing, brightness enhancement, filtering and denoising on the image in sequence to improve image quality and provide stable input for the subsequent process. The preset question-answer template is used to guide the DeepSeek-R1 large language model to automatically infer contextual background prompt words related to the target defect, further enriching the semantic expression of the image-text pair; and finally outputs a set of foreground text and background text phrases to guide subsequent image feature alignment and attention modeling. The support image and query image are then fed into the image encoder of the frozen pre-trained CLIP model to extract their multi-level visual features and multi-level attention weight matrices. Simultaneously, the generated foreground and background text are fed into the text encoder of the frozen pre-trained CLIP model to extract their corresponding semantic vectors (the foreground text embedding vector and each background text embedding vector). Semantic alignment between the visual and text features is established in the shared semantic space of the image and text encoders. Frozen pre-training directly uses the pre-trained weights; the image and text encoders of the CLIP model are not trained. Based on the query image's multi-level visual features, the foreground text embedding vector, and each background text embedding vector, the cosine similarity between them is calculated as the response strength. Grad-CAM is then used to perform gradient backpropagation on the classification layer that associates the query image with the text category, generating a saliency heatmap. Finally, the text response map is fused with the gradient heatmap to generate an initial activation map (coarse-grained activation map), which is then supervised and optimized using a cross-entropy loss function. Based on the multi-level visual features and multi-level attention weight matrix, a joint refinement strategy is used to obtain static and dynamic refinement matrices. The joint refinement strategy includes a static refinement branch and a dynamic refinement branch. The static refinement branch constructs a high-order optimization matrix based on the average attention weight matrix to enhance the consistency of the discriminative region; the dynamic refinement branch uses the activation map generated by the current model to convert it into a learnable pseudo-mask, and improves the diversity and stability of the activation map through an interactive learning strategy; finally, the static refinement activation map from the previous round of optimization is fused through matrix multiplication. and dynamically refined activation maps With coarse-grained activation map , generate the currently optimized static refined activation map and dynamically refined activation maps Based on static refinement activation map and dynamically refined activation maps Bootstrapping the few-shot segmentation model Segment the defective area in the query image to enhance the model's ability to focus on the defective area; and fuse the semantic features of the query image and the support image at multiple scales to obtain the defective area through the few-sample segmentation model. The final industrial defect segmentation prediction results are generated using a small sample size. The model training module trains the context-aware text enhancement model using all support-query pairs. The prediction module inputs the query image to be segmented, the corresponding support image, and the mask label into the trained context-aware text enhancement model. The output of the small sample segmentation model is used as the segmentation prediction result for the corresponding query image. The output segmentation prediction results, combined with evaluation indicators, can be used by quality inspection systems or manual review, thereby improving the accuracy and efficiency of image defect detection.
[0176] This system is proposed to address the problems of sample scarcity, labeling difficulties and insufficient model generalization ability in industrial defect detection tasks, especially in the realistic context where traditional few-shot learning methods ignore background semantic information, resulting in poor robustness of segmentation results. It uses the DeepSeek-R1 large language model to automatically generate foreground text and background text, introduces a foreground-background semantic linkage mechanism based on the image modality, jointly extracts visual and language features through the CLIP multimodal model, and uses the Grad-CAM method to generate an initial activation map (coarse-grained activation map). It further refines the activation area through a joint optimization strategy, guiding the segmentation model to focus on the target defect area more accurately. It has strong stability, versatility and engineering application prospects.
[0177] Regarding the specific limitations of the small sample industrial defect image segmentation system based on text background perception enhancement, please refer to the limitations of the small sample industrial defect image segmentation method based on text background perception enhancement above, which will not be repeated here. The various modules in the above-mentioned small sample industrial defect image segmentation system based on text background perception enhancement can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0178] The memory and processor are electrically connected, directly or indirectly, to enable data transmission or interaction. For example, these components may be electrically connected via one or more communication buses or signal lines. The memory stores a computer program executable on the processor. The processor executes the computer program stored in the memory to implement the small sample industrial defect image segmentation method based on text-context-aware enhancement in an embodiment of the present invention.
[0179] The memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM). The memory is used to store computer programs, and the processor executes the computer programs after receiving execution instructions.
[0180] Example 3:
[0181] A storage medium for small sample industrial defect image segmentation based on text background perception enhancement is used to store a computer program. When the computer program is executed by a processor, it implements the small sample industrial defect image segmentation method based on text background perception enhancement as described in Example 1.
[0182] When the computer program stored in the storage medium for small sample industrial defect image segmentation based on text background perception enhancement is executed by the processor, the small sample industrial defect image segmentation method based on text background perception enhancement as in Example 1 is implemented, thereby completing the small sample segmentation task of industrial defects. The processor may be an integrated circuit chip with data processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. The various methods disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The output segmentation prediction results combined with the evaluation indicators can be used by the quality inspection system or manual review, which is conducive to improving the detection accuracy and efficiency of image defects.
[0183] This storage medium is proposed to address the problems of sample scarcity, labeling difficulties and insufficient model generalization ability in industrial defect detection tasks, especially in the realistic context where traditional few-shot learning methods ignore background semantic information, resulting in poor robustness of segmentation results. It uses the DeepSeek-R1 large language model to automatically generate foreground text and background text, introduces a foreground-background semantic linkage mechanism based on the image modality, jointly extracts visual and language features through the CLIP multimodal model, and uses the Grad-CAM method to generate an initial activation map (coarse-grained activation map). It further refines the activation area through a joint optimization strategy, guiding the segmentation model to focus on the target defect area more accurately. It has strong stability, versatility and engineering application prospects.
[0184] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0185] The above-described embodiments merely represent specific and detailed examples of the present application and should not be construed as limiting the scope of the present application. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present application, and such modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A small sample industrial defect image segmentation method based on text background perception enhancement, characterized by: The steps include: S1. Construct corresponding support-query pairs based on industrial defect image sets ,in, For the query image, the support set , For the Support images, For the The mask labels of the supporting images, The number of supported images; S2. Build a text context-aware enhancement model to perform the following operations: S21, respectively obtaining a foreground text and multiple background texts of a defective area of each supporting image; S22, input the query image into the image encoder of the CLIP model, extract the multi-level visual features and the multi-level attention weight matrix, and input the foreground text and each background text of the support image of the corresponding category into the text encoder of the CLIP model, and obtain the foreground text embedding vector and each background text embedding vector; S23, generating a corresponding coarse-grained activation map using the Grad-CAM method based on the multi-level visual features of the query image, the foreground text embedding vector, and the background text embedding vectors; S24, obtaining a static refined activation map and a dynamic refined activation map according to the coarse-grained activation map, the multi-level visual features, and the multi-level attention weight matrix; S25. Input the query image, the corresponding support image, and the mask label into the few-shot segmentation model, and use the static refined activation map and the dynamic refined activation map as attention guidance signals; S3. Use all support-query pairs to train a text context-aware enhancement model; S4. Input the query image to be segmented, the corresponding support image and the mask label into the trained text background perception enhancement model, and obtain the output result of the few-shot segmentation model as the segmentation prediction result of the corresponding query image.
2. The small sample industrial defect image segmentation method based on text-background perception enhancement according to claim 1, characterized in that: The query image is an industrial defect image in the industrial defect image set, the support image is an industrial defect image with defect categories pre-labeled, and the mask label of the support image is a binary image or a multi-class labeled image corresponding to all pixels in the support image.
3. The small sample industrial defect image segmentation method based on text-background perception enhancement according to claim 1, characterized in that: The foreground text and multiple background texts of the defect area of each supporting image are generated by using the DeepSeek-R1 large language model according to the preset question-answering template; the multi-level visual features , For the layer visual features, the multi-level attention weight matrix , For the layer attention weight matrix, , is the total number of layers of the image encoder, To query the height of the image, The width of the query image.
4. The small sample industrial defect image segmentation method based on text-background perception enhancement according to claim 1, characterized in that: The Grad-CAM method is used to generate the corresponding coarse-grained activation map, as follows: S231, calculate the multi-level visual features The cosine similarity between the visual features of each spatial position of the layer visual feature and the foreground text embedding vector and the background text embedding vector, is the total number of layers of the image encoder; S232, normalizing each cosine similarity using a softmax function to form a corresponding probability vector; S233, the cross entropy loss between the probability vector and the true label is used as the coarse-grained activation map to generate the loss function, where the foreground of the true label is 1 and the background is 0; S234, calculate the coarse-grained activation map to generate the loss function for the The first layer of visual features Spatial position in the channel The partial derivative of the visual feature of the corresponding channel is the spatial position The gradient of the visual features, is the height direction index of the query image, is the width index of the query image, , , To query the height of the image, is the width of the query image; S235, the All gradients of each channel of the visual features of the layer are globally averaged and pooled to obtain the global gradient of the corresponding channel, and the global gradient of each channel is used to average the gradient of the first layer. The visual features of the corresponding channels of the layer visual features are weighted summed, and then the coarse-grained activation map is generated through the ReLU activation function.
5. The small sample industrial defect image segmentation method based on text background perception enhancement according to claim 1 is characterized in that: According to the coarse-grained activation map, the multi-level visual features and the multi-level attention weight matrix, the static refined activation map and the dynamic refined activation map are obtained as follows: S241. Obtaining a static refinement matrix and a dynamic refinement matrix based on a joint refinement strategy according to the multi-level visual features and the multi-level attention weight matrix; S242. Perform matrix multiplication operations on the coarse-grained activation map with the static refinement matrix and the dynamic refinement matrix, respectively, to obtain a static refinement activation map and a dynamic refinement activation map, respectively.
6. The small sample industrial defect image segmentation method based on text-background perception enhancement according to claim 5, characterized in that: After obtaining the static refined activation map and the dynamic refined activation map based on the coarse-grained activation map, the multi-level visual features, and the multi-level attention weight matrix, a pseudo label is generated from the dynamic refined activation map according to a preset threshold, that is, when the activation value of each spatial position in the dynamic refined activation map is greater than or equal to the preset threshold, it is considered to be a foreground area of the pseudo label; otherwise, it is considered to be a background area of the pseudo label; The joint refinement strategy includes a static refinement strategy and a dynamic refinement strategy, wherein: The static refinement strategy performs the following operations: The attention weight matrix of each layer in the multi-level attention weight matrix is averaged and the normalized matrix is obtained using the Sinkhorn algorithm; Perform matrix multiplication operation on the normalized matrix and the transpose of the normalized matrix to obtain a high-order optimized matrix; Take the larger value of the high-order optimization matrix and the normalized matrix at each spatial position to generate a static refinement matrix; The dynamic refinement strategy performs the following operations: Concatenate all the visual features of the multi-level visual features in the channel dimension and perform convolution operation Obtaining intermediate layer features; Perform a shape reshaping operation on the intermediate layer features to form a reshaped feature, and then perform a matrix multiplication operation on the reshaped feature and the transpose of the reshaped feature to form an initial dynamic refinement matrix; Calculate the difference between the attention weight matrix of each layer in the multi-level attention weight matrix and the corresponding element in the initial dynamic refinement matrix, and add the differences of all elements in the corresponding layer to form the fusion feature of the corresponding layer; The fusion features of each layer are averaged to obtain the average matrix, and the attention weight matrix with higher than the average matrix in the multi-level attention weight matrix is selected as the middle layer attention matrix; The selected middle layer attention matrix is averaged over the channels to obtain the middle layer optimization matrix; The initial dynamic refinement matrix is multiplied pixel by pixel with the intermediate layer optimization matrix to obtain the dynamic refinement matrix of the current training round, and the cross entropy loss between the dynamic refinement matrix of the current training round and the pseudo label is used as the loss function of the dynamic refinement matrix of the next training round.
7. The small sample industrial defect image segmentation method based on text background perception enhancement according to claim 1 is characterized in that: The segmentation loss of the few-shot segmentation model is the cross entropy loss between the segmentation prediction result of the query image and the true mask label.
8. The small sample industrial defect image segmentation method based on text background perception enhancement according to claim 1 is characterized in that: The few-shot segmentation model is a PFENet model or a HDMNet model.
9. A small sample industrial defect image segmentation system based on text background perception enhancement, characterized by: It includes data acquisition module, model building module, model training module and prediction module, among which: The data acquisition module is used to construct corresponding support-query pairs based on the industrial defect image set. ,in, For the query image, the support set , For the Support images, For the The mask labels of the supporting images, The number of supported images; The model building module is used to build a text background perception enhancement model to perform the following operations: respectively obtaining a foreground text and a plurality of background texts of a defect area of each supporting image; The query image is input into the image encoder of the CLIP model, and multi-level visual features and multi-level attention weight matrices are extracted accordingly. The foreground text and background text of the support image of the corresponding category are input into the text encoder of the CLIP model, and the foreground text embedding vector and background text embedding vector are obtained accordingly. According to the multi-level visual features of the query image, the foreground text embedding vector and the background text embedding vector, the Grad-CAM method is used to generate the corresponding coarse-grained activation map; Obtain static refined activation maps and dynamic refined activation maps based on the coarse-grained activation map, multi-level visual features, and multi-level attention weight matrix; The query image, the corresponding support image, and the mask label are fed into the few-shot segmentation model, and the static refined activation map and the dynamic refined activation map are used as attention guidance signals. The model training module is used to train the text context perception enhancement model using all support-query pairs; The prediction module is used to input the query image to be segmented, the corresponding support image and the mask label into the trained text background perception enhancement model, and obtain the output result of the few-shot segmentation model as the segmentation prediction result of the corresponding query image.
10. A storage medium for small sample industrial defect image segmentation based on text-background perception enhancement, used to store a computer program, characterized by: When the computer program is executed by a processor, the method for segmenting small sample industrial defect images based on text background perception enhancement is implemented as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Surface defect small sample detection method based on contrast learning
CN119090866A
Zero-sample industrial defect detection method and equipment based on text guidance and medium
CN119784670A
Less-annotation remote sensing image semantic segmentation method based on visual text guidance
CN120070895A
Cited By
Weak supervision image defect segmentation method and system based on text guidance
CN120807555A
Multi-modal driving and dynamic optimization power station inspection image segmentation method and multi-modal driving and dynamic optimization power station inspection image segmentation system
CN122289698A
Power station inspection image segmentation method and system based on multi-modal driving and dynamic optimization
CN122289698B