Object surface defect segmentation method based on small samples

By constructing a defect segmentation network for small sample learning and using the query prior mask generator and feature interaction module to compensate for information loss, the problems of intra-class differences and multi-scale changes in metal surface defect segmentation are solved, efficient and accurate defect segmentation is achieved, and the generalization ability of the model is improved.

CN120689623APending Publication Date: 2025-09-23WUHAN TEXTILE UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510808362.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing metal surface defect segmentation methods based on deep convolutional neural networks rely on large-scale high-quality datasets, are difficult to adapt to complex industrial environments, and have insufficient generalization capabilities for new categories. In particular, the segmentation performance deteriorates when there are significant intra-class differences between metal defect samples, global pooling causes spatial information loss, and the internal scale changes of a single defect image are complex.

Method used

A defect segmentation network is constructed using a small sample learning-based method, including a query prior mask generator, a prior-guided bidirectional feature interaction module, a prototype-loss compensation module, and a context-aware attention-guided feature aggregation decoding module. By generating query prior masks, enhancing feature interactions, compensating for lost information, and capturing multi-scale contextual information, efficient and accurate defect segmentation is achieved.

Benefits of technology

Under the condition of a small number of samples, the accuracy and robustness of metal surface defect segmentation are improved, intra-class differences and multi-scale changes are effectively handled, computing resource consumption and mismatching are reduced, and the generalization ability of the model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689623A_ABST
    Figure CN120689623A_ABST
Patent Text Reader

Abstract

The invention discloses an object surface defect segmentation method based on small samples, and the method specifically comprises the steps: collecting images of different types of object surface defects and corresponding segmentation masks, constructing a small sample data set, and dividing the small sample data set into a training set and a test set; constructing training and testing tasks, and respectively extracting a corresponding support set and a query set for each task; constructing and training a defect segmentation network, wherein the defect segmentation network comprises a query prior mask generator, a prior-guided bidirectional feature interaction module, a prototype-loss compensation module and a context-aware attention-guided feature aggregation decoding module; and after training is completed, inputting the support set and the to-be-tested query image into the converged defect segmentation network, and outputting a query image prediction result. Experiments prove that the method realizes remarkable performance improvement on an FSSD-12 data set, and shows that the method has obvious advantages in a small sample defect segmentation task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method for segmenting object surface defects based on small samples. Background Art

[0002] In modern industrial production processes, accurate and efficient inspection of product surface quality is crucial for ensuring product quality, enhancing brand reputation, and reducing production costs. Especially in fields such as metal processing, automotive manufacturing, aerospace, and electronics production, minute surface defects on metal parts, such as scratches, wear, inclusions, oil stains, and cracks, can seriously impact product performance, safety, and service life.

[0003] Traditional methods for metal surface defect segmentation rely primarily on manual visual inspection, which suffers from low efficiency, susceptibility to subjective factors, difficulty detecting minor defects, and high labor costs. With the development of computer vision technology, automated defect detection systems based on machine vision have emerged. However, early methods relied on manually designed image processing algorithms and traditional machine learning models, resulting in poor robustness and difficulty adapting to complex industrial environments. In recent years, deep learning, especially deep convolutional neural networks, has made breakthrough progress in the field of image processing. Fully supervised defect segmentation methods based on deep convolutional neural networks can automatically learn defect features and achieve segmentation accuracy far exceeding traditional methods on specific datasets. Network structures such as U-Net, FCN, DeepLab, and their variants have been widely used.

[0004] Although defect segmentation methods based on deep convolutional neural networks have performed well, their success is heavily dependent on large-scale, high-quality annotated datasets. With technological advancements and the improvement of product quality in real-world industrial scenarios, it is becoming increasingly difficult to obtain large-scale, high-quality datasets. Not only is pixel-level precise annotation time-consuming and labor-intensive, resulting in high annotation costs, but high-quality production control also makes defect samples relatively scarce, making it difficult to collect sufficient training data, especially for rare defect categories. What's worse, continuous adjustments to the production process may also lead to the emergence of new defect types, and methods based on deep convolutional neural networks often lack the ability to generalize to new categories, making it difficult to adapt to rapid changes in production lines. Therefore, how to achieve efficient and accurate industrial defect segmentation when annotated data is limited or even absent has become an urgent problem to be solved.

[0005] In recent years, small-sample learning technology has been introduced to the field of defect segmentation, forming a defect segmentation method based on small-sample learning. Its goal is to enable the model to perform pixel-level segmentation on new, unseen samples by learning only a small number of labeled samples. This greatly reduces the data requirements for model training and improves model generalization.

[0006] Existing methods for small-sample defect segmentation typically follow a meta-learning paradigm, learning transferable representations from base classes and applying them to new categories. In both the training and testing phases, the samples used are divided into a support set and a query set. Given limited data, the model must effectively utilize the information contained in the support images (derived from the support set) to accurately segment the query image (derived from the query set). Among these, methods based on prototype networks are an important class of solutions. By encoding support set images into compact feature prototypes, they aim to capture the essential characteristics of defects in that category and thereby guide the segmentation of query images. However, existing methods still face the following challenges:

[0007] (1) There are significant intra-class differences between metal defect samples, making it difficult to effectively generate prototypes with good discrimination capabilities. The morphology of metal defects is extremely diverse, and defects of the same category may have significant differences in texture, contrast, location, and lighting conditions during imaging. This difference makes it difficult for the prototype learned from a very small number of support samples to fully capture the essential characteristics of the category, resulting in a sharp drop in performance when segmenting query samples with large morphological differences. Although existing methods have attempted to solve this problem, they often have defects such as high computational resource consumption, easy mismatching, insufficient feature interaction, and the need to introduce additional defect-free images for auxiliary training.

[0008] (2) Global pooling causes spatial information loss and limits prototype representation capabilities. The global pooling operation during prototype generation inevitably causes spatial information loss. For defect images with good features, this loss may be limited to partial edge information; however, for samples with complex characteristics, global pooling may lose a large amount of key information, seriously hindering effective prototype representation and subsequent segmentation.

[0009] (3) The scale variation within a single defect image is complex, making multi-scale feature fusion difficult. Within a single defect image, the size of the defect may range from tiny point damage to large-scale regional damage. After the prototype and query image feature fusion, if this multi-scale information cannot be well processed, it will make it difficult for the model to accurately identify and segment defects of different scales, ultimately affecting the segmentation performance.

[0010] In view of this, it is necessary to provide a small sample segmentation method for object surface defects based on small samples to solve the above-mentioned defects of the prior art. Summary of the Invention

[0011] In order to accurately locate and segment defects using a small number of surface defect sample images, the present invention proposes a surface defect segmentation method based on a small sample, which specifically includes the following steps:

[0012] Step S1: collect images of surface defects of different categories of objects and their corresponding segmentation masks, construct a small sample data set and divide it into a training set and a test set;

[0013] Step S2: construct training and test tasks, and extract corresponding support sets and query sets for each task;

[0014] Step S3, constructing and training a defect segmentation network, wherein the defect segmentation network includes a query prior mask generator, a prior-guided bidirectional feature interaction module, a prototype-loss compensation module, and a context-aware attention-guided feature aggregation decoding module;

[0015] The query priori mask generation module is used to generate a query image priori mask map;

[0016] The prior-guided bidirectional feature interaction module is used to obtain enhanced support features and enhanced query features, and generate support prototypes;

[0017] The prototype-loss compensation module predicts the support image using the support prototype to obtain the predicted segmentation result of the support image and generates a compensation prototype;

[0018] The context-aware attention-guided feature aggregation decoding module is used to fuse the enhanced query features with the supporting prototype and the compensating prototype to ultimately generate a pixel-level defect segmentation prediction result for the query image;

[0019] Step S4: After the training is completed, the support set and the query image to be tested are input into the converged defect segmentation network, and the query image prediction result is output.

[0020] Furthermore, the defect segmentation network processes as follows: given a pair of support samples and query image , using two weight-sharing feature extraction networks to extract multi-level support features and query features ,in, and They are the support image in the support set and its corresponding segmentation mask, and They represent the feature extractor The feature map output by the layer; for subsequent processing, the third and fourth level features are spliced ​​and fused to generate , ; Using the fourth level features to generate the query image prior mask and the query prior mask generator and After filtering respectively, the prior-guided bidirectional feature interaction module is used to ensure sufficient feature interaction between the two features and generate enhanced support features respectively. and enhanced query features ; Then, through global average pooling, using Generate support prototype At this time, the prototype-loss compensation module uses the support prototype to predict the support image and obtain the predicted segmentation result of the support image , thereby detecting the spatial information lost during the generation of the main prototype and generating a compensation prototype , and provide auxiliary loss; finally, the context-aware attention-guided feature aggregation decoding module is integrated with and the prototype from the support branch Finally, the features are aggregated and decoded to predict the query image , generate segmentation results .

[0021] Furthermore, a feature extraction network pre-trained on the ImageNet dataset with frozen parameters is used to process the support image and query image in the current training or test task, respectively, and the third-level features and fourth-level features are preferably extracted for subsequent feature processing; wherein, the third-level support and query features are respectively denoted as and , the support and query features of the fourth level are respectively recorded as and ;

[0022] The generation process of fusion features is as follows:

[0023]

[0024]

[0025] in, and Represent the calculated support fusion features and query fusion features respectively, Indicates concatenation of the two input feature maps in the channel dimension. Represents a The convolution layer receives the channel-joined feature map as input and adjusts its channel number to the preset fusion channel number through convolution operation.

[0026] Furthermore, the specific implementation process of querying the priori mask generator in step S3 includes:

[0027] Generate foreground prototype and background support prototype:

[0028]

[0029]

[0030] in, and Indicates the fourth level of support and query features, represents the masked average pooling operation, is the spatial resolution of the fourth-level features, and are the calculated foreground and background prototype vectors, and They are respectively Downsampled foreground and background support masks with consistent spatial resolution; and ,in Indicates that the original image mask is adjusted to the same spatial resolution as the fourth-level feature through bilinear interpolation downsampling operation. To support mask; is the index of all spatial locations, represents element-wise multiplication;

[0031] Measuring Level 4 Support Characteristics The similarity between the features of each spatial position in and the support prototype calculated above:

[0032]

[0033]

[0034] in, and Represent the similarity graphs of the support vector to the foreground and background of the query image, respectively. is the index of all spatial positions, symbol Represents a vector norm, is a very small positive number added to the denominator to ensure numerical stability, and Respectively represent taking the minimum value and taking the maximum value;

[0035] Normalize the two original similarity graphs:

[0036]

[0037]

[0038] in, and denote the initial foreground and background prior masks of the query image, respectively;

[0039] Perform the blur region identification and removal steps. By calculating the Hadamard product of the two normalized similarity maps, we can identify the regions composed of blurred pixels. These regions are then removed from the foreground prior mask to obtain the purified prior mask. Finally, the query prior mask is normalized again to obtain the final required query prior mask:

[0040]

[0041]

[0042]

[0043] in, Indicates the fuzzy area, Represents the purified mask, where “−” represents the removal operation. Indicates that the bilinear interpolation upsampling operation will Adjust to the same spatial resolution as the query image to obtain the final query prior mask.

[0044] Furthermore, the specific implementation process of the prior-guided bidirectional feature interaction module in step S3 includes:

[0045] Perform feature masking operations to reduce the interference of complex background. The two feature masking operations can be expressed as:

[0046]

[0047]

[0048] in, Represents the supported mask The support features after filtering, Represents the query prior mask The filtered query features, and They are respectively and Downsampling mask with consistent spatial resolution;

[0049] The masked support features and the masked query features obtained after preprocessing are input into the bidirectional cross attention unit to establish pixel-level correspondence and achieve mutual feature enhancement. The process uses the following formula to enhance the query and support features:

[0050]

[0051]

[0052] in, and Represent the support features and query features that are initially enhanced after interaction, 、 、 Both represent convolution, represents matrix multiplication, are the indices of all possible positions, Represents a feature map After convolution mapping at position The eigenvectors on Represents another feature map passing through After convolution mapping at position The eigenvectors on , Indicates that the similarity between two feature vectors is calculated by dot product operation. is the number of positions in the input feature map that participate in the attention calculation;

[0053] In order to further utilize the global context information of the supporting samples to optimize the features after interaction, a squeezing and co-excitation mechanism is introduced to adaptively reweight the channel dimension of the initially enhanced features. The weight calculation formula can be expressed as:

[0054]

[0055] in, is the final calculated channel attention weight vector, represents the global average pooling operation, represents a multi-layer perceptron network used to learn nonlinear inter-channel dependencies from global channel descriptors and generate attention weights;

[0056] Finally, the calculated channel attention weights Support features applied to preliminary enhancements and query features , and then weighted by multiplying each channel to obtain the final enhanced support feature and enhanced query features ; The calculation formula is as follows:

[0057]

[0058]

[0059] Among them, among them, and Represent the final enhanced support features and query features, As the index of all spatial locations, As the index of the feature channel, Indicates the The attention weight of each channel, Represents simple numeric multiplication.

[0060] Furthermore, the specific implementation process of the prototype-loss compensation module generating the compensated prototype in step S3 includes:

[0061] The enhanced support features output by the prior-guided bidirectional feature interaction module generate the main support prototype, and its calculation formula is as follows:

[0062]

[0063] in, Represents the calculated main support prototype;

[0064] The main support prototype By copying and expanding the spatial dimension, its spatial resolution is consistent with the original support fusion feature Then, the original support is integrated with the special With two copied and extended main prototypes The calculation formula for splicing in the channel dimension is:

[0065]

[0066] in, are input features used to support image prediction, Indicates that the vector is expanded in the spatial dimension;

[0067] Perform a self-prediction on the support samples to discover information that the main prototype may have overlooked and generate a compensating prototype:

[0068]

[0069] in is the pixel-level segmentation prediction mask of the support image under the guidance of the main prototype, The decoder is a context-aware attention-guided feature aggregation decoding module;

[0070] Based on the support image prediction mask and the true support foreground mask, all correctly predicted foreground pixels are calculated using the following formula:

[0071]

[0072] in Gather all correctly predicted foreground pixels, representing the correctly segmented foreground area. " represents the logical AND operation;

[0073] Remove the correctly segmented foreground regions from the true support foreground mask to get all incorrectly predicted foreground pixels:

[0074]

[0075] in, represents the foreground information lost during the prototype generation process, " represents the removal operation;

[0076] The surrounding background information of the supporting image is used to supplement the defect. The complete compensation mask is defined as follows:

[0077]

[0078] The compensation prototype is obtained by mask average pooling using the compensation mask and fusion support features. The specific formula is as follows:

[0079]

[0080] in, is After downsampling, Indicates that the bilinear interpolation downsampling operation will be Adjust to the characteristics The same spatial resolution, is the calculated compensation prototype.

[0081] Furthermore, the implementation process of step S3 specifically includes:

[0082] Enhanced query features , Main Support Prototype and compensation prototypes To perform fusion, take as input:

[0083]

[0084] use Convolution operation generates compressed features ,in ; In order to capture multi-scale context information, we further Input into the void space pooling pyramid module to form , this process is written as:

[0085]

[0086] in, Contains a global average pooling operation, a Convolutional layer and a dilation operation; Represents multiple parallel hole convolution layers, subscript is the void ratio, Represents the ReLU activation function;

[0087] In order to further enhance the feature representation, an attention-guided residual block with two branches is adopted, namely the convolution branch and the attention branch. The output of the attention-guided residual block of the two branches is expressed as:

[0088]

[0089] The final predicted segmentation mask for the query image is obtained through a fully convolutional decoder:

[0090]

[0091] in, is the final prediction result of the query image, Represents a fully convolutional decoder, consisting of a Convolutional layer, a ReLU function activation layer, a Convolutional layers, represents a bilinear interpolation upsampling operation used to adjust the predicted mask to the same spatial resolution as the query image, Used to obtain the segmentation mask, which means using this function to determine the category to which each pixel belongs from the output of the model, thereby generating the final segmentation mask.

[0092] Furthermore, the loss function calculation formula used in training the defect segmentation network in step S3 is specifically as follows:

[0093]

[0094]

[0095]

[0096] in, and denote the cross entropy loss in query image segmentation prediction and support image segmentation prediction, respectively. From Prototype-Loss Compensation Module, Auxiliary loss The proportion of and Represent the segmentation prediction probability maps corresponding to the support image and the query image respectively, and , represents the number of pixels in the probability map, is the index of the pixel, represents the number of samples in the support set, is the index of the sample.

[0097] The present invention also provides a computer program product comprising computer program instructions, which are stored on a computer-readable storage medium. When the instructions are executed by a processor, the method for segmenting surface defects of an object based on a small sample as described in the above technical solution is implemented.

[0098] The present invention also provides an electronic device, comprising a memory and a processor communicatively connected to the memory, wherein the memory stores computer program instructions, and when the processor executes the computer program instructions, it implements a method for segmenting surface defects of an object based on a small sample as described in the above technical solution.

[0099] The working principle and beneficial effect of this technical solution lies in proposing a new small-sample defect segmentation network to solve the challenging few-shot metal surface defect segmentation problem, especially in dealing with significant intra-class differences, complex defect scale variations, and spatial information loss. It includes four core components: a query prior mask generation module, a prior-guided bidirectional feature interaction module, a prototype-loss compensation module, and a context-aware attention-guided feature aggregation and decoding module.

[0100] (1) The query prior mask generation module generates a query image prior mask map, providing a preliminary, low-cost positioning indication of potential target regions in the query image. This prior mask can effectively guide the subsequent prior-guided bidirectional feature interaction module to focus more on regions in the query image that are related to the target categories in the support set, reducing the probability of false matches caused by background interference or intra-class changes, and improving the efficiency and robustness of feature interaction.

[0101] (2) The prior-guided bidirectional feature interaction module guides the bidirectional feature interaction based on cross-attention through an innovative training-independent query prior mask generation mechanism. Its role is to alleviate the problem of significant intra-class differences between defective samples, and solve the defects of previous methods such as high computational resource consumption, easy mismatching, insufficient feature interaction, and the need to introduce additional defect-free image auxiliary training, so as to achieve more sufficient and robust information exchange between the support set and query set features, and help generate more discriminative support prototypes.

[0102] (3) The prototype-loss compensation module identifies the information lost during the main prototype generation process by making predictions in the support branch, and generates a compensatory prototype based on this to make up for the lost spatial details and difficult-to-distinguish sample information. At the same time, an auxiliary loss is introduced to strengthen supervision. Its purpose is to alleviate the spatial information loss caused by global pooling and improve the integrity of the prototype representation.

[0103] (4) The context-aware attention-guided feature aggregation decoding module uses a dilated spatial convolutional pooling pyramid to capture multi-scale contexts and introduces coordinate attention to achieve position-sensitive and direction-aware attention guidance, aiming to effectively aggregate feature information at different scales, thereby accurately processing the complex scale changes in defect images. BRIEF DESCRIPTION OF THE DRAWINGS

[0104] Figure 1 This is a flow chart of a method for segmenting object surface defects based on small samples provided by the present invention.

[0105] Figure 2 The surface defect segmentation network model for small sample objects constructed in the present invention is demonstrated from a macroscopic perspective.

[0106] Figure 3 Schematic diagram of the structure of the query priori mask generator in the present invention.

[0107] Figure 4 Schematic diagram of the structure of the a priori-guided bidirectional feature interaction module in the present invention.

[0108] Figure 5 Schematic diagram of the structure of the prototype-loss compensation module in the present invention.

[0109] Figure 6 Schematic diagram of the structure of the context-aware attention-guided feature aggregation decoding module in the present invention. DETAILED DESCRIPTION

[0110] The following will describe in more detail the objectives, technical solutions and advantages of the present invention with reference to the accompanying drawings and specific embodiments.

[0111] The present invention provides a new method for segmenting surface defects of objects based on small samples, such as Figure 1 As shown, the method includes the following steps S1-S4:

[0112] Step S1, collecting or acquiring images containing different categories of surface defects of metal workpieces and their corresponding segmentation masks, constructing a dataset for small sample defect segmentation, and dividing the dataset into a training set and a test set;

[0113] In step S1, the purpose of dividing the dataset into training and test sets is only to perform cross-validation, which is an effective way to make full use of limited data to improve the generalization ability of the model. Due to the limited size of the dataset for small sample defect segmentation, directly dividing it into fixed training, validation, and test sets will further reduce the training data and affect the learning effect of the model. Through cross-validation, the training set can be divided into multiple subsets, which are used as test sets in turn, and the remaining subsets are used as training sets. The model is trained and tested multiple times, thereby making full use of all data for learning and evaluation. It is particularly important that the training set and the test set not only contain completely different defect samples, but also belong to completely different types. This requires the model to have strong generalization capabilities and be able to identify and segment defect types that have not been seen in the training set.

[0114] Step S2: construct training and test tasks, and extract corresponding support sets and query sets for each task;

[0115] Specifically, for each episode in the training phase, first, a defect category is randomly selected from the training set. Then, sample pairs as support set And randomly select a sample pair as the query set .in, and They are respectively the support concentration images and their corresponding segmentation masks, and are the images in the query set and their corresponding segmentation masks. The support set contains a small number of images with segmentation masks, representing the defect categories that need to be learned in this round. The query set contains other images of the defect category. The model needs to predict the segmentation masks of the defects in the query set based on the information in the support set. This process is repeated multiple times, each time randomly selecting a different defect category and constructing a new support set and query set, so that the model can learn the common features of various defects and ultimately improve its ability to segment unknown defects on the test set. For each round in the testing phase, the extraction process of the support set and query set is the same as that in the training phase. The difference is that in the training phase, the query image mask is used to supervise the model to optimize its own parameter learning and performance evaluation, while in the testing phase, the model parameters are frozen and the query image mask is only used to evaluate the model performance. Preferably, during the data and feature processing, whether in the training phase or the testing phase, only the support image mask can be used, and the query mask is not allowed to be used unless it is used to calculate the loss function in the training phase or for performance evaluation in both phases.

[0116] Step S3, constructing and training a defect segmentation network, wherein the defect segmentation network includes a query prior mask generator, a prior-guided bidirectional feature interaction module, a prototype-loss compensation module, and a context-aware attention-guided feature aggregation decoding module;

[0117] In order to more clearly explain the specific implementation details of the small sample metal surface defect segmentation network constructed by the present invention, the present invention divides the above steps into six parts for detailed description: model process overview, multi-level feature extraction and predetermined level feature fusion, query prior mask generation, prior-guided bidirectional feature interaction, prototype-loss compensation, and context-aware attention-guided feature aggregation decoding, and loss function calculation. Specifically, it includes the following steps:

[0118] (1) Overview of the model process

[0119] Figure 2 The small sample metal surface defect segmentation network model constructed in the present invention is presented from a macro perspective. The figure shows the overall feature flow process of image data from input to output, aiming to help readers understand how image data and feature information are transmitted in the model. Specifically, the network model constructed in the present invention is mainly composed of four key components: query prior mask generation, prior-guided bidirectional feature interaction module, prototype-loss compensation module, and context-aware attention-guided feature aggregation decoding module. In the 1-shot (support set contains only a pair of support images and their masks) setting, given a pair of support samples and query image , using two weight-sharing feature extraction networks to extract multi-level support features and query features .in, and They represent the feature extractor For subsequent processing, the third and fourth level features are concatenated and fused to generate , . Using the fourth level features to generate the query image prior mask and the query prior mask generator and After filtering respectively, the prior-guided bidirectional feature interaction module is used to ensure sufficient feature interaction between the two features and generate enhanced support features respectively. and enhanced query features Then, we use Global Average Pooling (GAP) to Generate support prototype At this point, the prototype-loss compensation module uses the support prototype to predict the support image and obtain the predicted segmentation result of the support image , thereby detecting the spatial information lost during the generation of the main prototype and generating a compensation prototype , and provides auxiliary loss. Finally, the context-aware attention-guided feature aggregation decoding module is integrated with and the prototype from the support branch Finally, the features are aggregated and decoded to predict the query image , generate segmentation results The above is a brief overview of the overall architecture and workflow of the neural network model of the present invention. The detailed processing process and specific details of each module will be elaborated in the subsequent content.

[0120] (2) Multi-level feature extraction and pre-defined level feature fusion

[0121] Regarding multi-level feature extraction, the present invention utilizes a feature extraction network (e.g., a mature convolutional neural network architecture such as VGG-16 or ResNet-50) pre-trained on the ImageNet dataset with frozen parameters to process the support image and query image in the current training (or testing) task separately. The feature extraction network preferably employs a weight sharing mechanism, using the same set of network parameters to process both the support image and the query image to ensure consistency in the feature space. The feature extraction network is configured to output feature maps at multiple different levels. These feature maps typically correspond to the outputs of feature layers at different depths within the feature extraction network (e.g., for a ResNet architecture, these are Layers 1 through 5, representing the five main feature extraction stages within the network). Typically, first-level features (Layer 1 output) and second-level features (Layer 2 output) are referred to as low-level features. They focus on capturing basic image information such as edges, color, and texture, but are relatively weak in expressing semantic information. The third-level (Layer 3 output) and fourth-level (Layer 4 output) features are called mid-level features and can express the local structure of the image and object component information. As an effective abstraction of low-level features, they contain relatively rich semantic information while also well preserving spatial details. The fifth-level (Layer 5 output) features are called high-level features. These features can better express the categorical attributes of the image, but after multiple downsampling and nonlinear transformations, they lose a lot of spatial information, making it difficult to accurately locate the position and boundaries of the target object.

[0122] Regarding the fusion of predetermined-level features, the present invention effectively integrates features from different depths of the feature extraction network to generate a unified feature representation that can retain sufficient spatial resolution for precise positioning and capture high-level semantic context for accurate identification. This fused feature will serve as the input of subsequent high-level processing modules (such as the prior-guided bidirectional feature interaction module proposed in the present invention). The feature fusion process is performed independently on the features of the support samples and the features of the query samples. Specifically, in view of the problems of insufficient semantic expression ability of low-level features and more loss of spatial information of high-level features, the present invention preferably takes intermediate features (third and fourth-level features) for subsequent feature processing. Among them, the third-level support and query features are respectively denoted as and , and the fourth-level support and query features are respectively recorded as and Before fusion, ensure that the feature maps of these selected levels have the same spatial resolution (if the original output is inconsistent, adjust it by upsampling). The generation process of fused features is as follows:

[0123]

[0124]

[0125] in, and Represent the calculated support fusion features and query fusion features respectively, Indicates concatenation of the two input feature maps in the channel dimension. Represents a The convolution layer receives the channel-joined feature map as input and adjusts the number of channels to a preset number of fusion channels through convolution operation. The preset number of fusion channels in the present invention is 256.

[0126] (3) Query prior mask map generation

[0127] Figure 3This is a schematic diagram of the structure of the query prior mask generator in the present invention. It aims to generate a prior mask map for the query image. The role of this mask map is to preliminarily indicate the areas in the query image where target defects may exist. In subsequent feature interactions, the interference of background noise in the query image is minimized as much as possible, guiding the model to focus more on potential foreground targets. More importantly, this generation process does not require the introduction of additional learning parameters. The prior mask maps generated by previous prior mask map generators contain some areas composed of pixels of poor quality. These pixels can be divided into two categories: blurry area pixels and incorrectly activated noise pixels. Among them, blurry area pixels are caused by the convolution kernel performing convolution operations on the boundary of the target defect area during the feature extraction process, which may mix foreground and background information at the same time, often causing the estimated boundary of the defect in the query prior mask map to be over-expanded. Incorrectly activated noise pixels are caused by foreground pixel values ​​of the support image mistakenly generating high responses to background pixel values ​​of the query image. They usually appear at the same time as blurry area pixels, making it more difficult to generate high-quality prior mask maps. Based on the above observations, this paper proposes a novel query prior mask map generator, which aims to reduce the occurrence of these two types of pixels and provide more accurate guidance information.

[0128] Considering that prototype features are usually more robust to outliers and help reduce the impact of false activation noise, the first step of querying the prior mask map generation in the present invention is to use the aforementioned fourth-level support features and the corresponding support mask , the foreground support prototypes are calculated separately through the Masked Average Pooling (MAP) operation and background support prototypes , the specific calculation formula of this process is:

[0129]

[0130]

[0131] in, represents the masked average pooling operation, is the spatial resolution of the fourth-level feature, representing the total number of spatial locations in the feature map. and are the calculated foreground and background prototype vectors, and They are respectively Downsampled foreground and background support masks with consistent spatial resolution. Note that and ,in It means that the original image mask is adjusted to the same spatial resolution as the fourth-level feature through bilinear interpolation downsampling operation. is the index of all spatial locations, Represents element-wise multiplication.

[0132] Next, we need to measure the fourth level support characteristics The similarity between the feature of each spatial position in the query feature and the support prototype calculated above. For each spatial position in the query feature, calculate its similarity with the foreground prototype and background prototype The calculation formula of this process is:

[0133]

[0134]

[0135] in, and Represents the similarity graph of the support vector to the foreground and background of the query image, respectively, and the symbols Represents a vector norm, is a very small positive number (e.g., 1e-8) added to the denominator to ensure numerical stability, and Respectively represent the minimum and maximum values.

[0136] Subsequently, in order to unify the numerical range and facilitate subsequent operations, the two original similarity graphs need to be normalized. The present invention adopts the minimum-maximum normalization method to process the two similarity graphs independently. and The calculation formula is:

[0137]

[0138]

[0139] in, and denote the initial foreground and background prior masks of the query image, respectively.

[0140] Next, the blurred region identification and removal step is performed to process pixels with high foreground and background responses. Specifically, the Hadamard product (element-wise multiplication) of the two normalized similarity maps is calculated to identify regions composed of blurred pixels. These regions are then removed from the foreground prior mask to obtain a purified prior mask map. Finally, the map is normalized again to obtain the final query prior mask map:

[0141]

[0142]

[0143]

[0144] in, Indicates the fuzzy area, Represents the purified mask, where “−” represents the removal operation. Indicates that the bilinear interpolation upsampling operation will Adjust to the same spatial resolution as the query image to obtain the final query prior mask.

[0145] (4) Prior-guided bidirectional feature interaction

[0146] Specifically, the present invention implements this process through a priori guided two-way feature interaction module, such as Figure 4 As shown, this module utilizes the fusion support features mentioned above , query features , query the prior mask , and support for image masks , obtaining more discriminative and better aligned enhanced support features and enhanced query features.

[0147] First, a feature mask operation is performed. In order to reduce the interference of complex background, a support mask is used Support fusion features Filtering is performed to highlight potential foreground areas and suppress irrelevant background information; at the same time, the query prior mask is used Fusion features for queries Filtering is performed to suppress the presence of background pixels as much as possible. The two feature mask operations can be expressed as:

[0148]

[0149]

[0150] in, and They are respectively and Downsampling mask with consistent spatial resolution, Represents the supported mask The support features after filtering, Represents the query prior mask Filtered query features

[0151] Next, the mask support features obtained after preprocessing and mask query features The input is fed into a bidirectional cross-attention unit to establish pixel-level correspondences and achieve mutual feature enhancement. This unit uses a non-local block to build correlations between each element and all other elements. Each element can associate itself with all data elements, forming a long-range dependency. This process enhances the query and support features using the following formula:

[0152]

[0153]

[0154] in, and Represent the support features and query features that are initially enhanced after interaction, 、 、 Both represent Convolution is used to map input features to a unified low-dimensional space to reduce computational complexity and improve the ability to model the correlation between features, so as to better calculate correlation or attention weights. represents matrix multiplication, are the indices of all possible positions, Represents a feature map After convolution mapping at position The eigenvectors on Represents another feature map passing through After convolution mapping at position The eigenvectors on , Indicates that the similarity between two feature vectors is calculated by dot product operation. In order to prevent the value of the dot product result from being too large, which may cause gradient disappearance or instability in subsequent operations, the present invention adopts the strategy of scaled dot product, that is, multiplying the dot product result by , is the number of locations in the input feature map that participate in the attention computation. The above formula for enhancing query and support features implies some tensor operations, for example, After convolution mapping, the feature map is usually flattened to facilitate the calculation of dot products and subsequent attention weights. In order to meet the dimensional requirements of matrix multiplication, one of the feature maps may also need to be transposed. After the attention weight is calculated, it is necessary to multiply it with the feature map and reshape the result back to the shape of the original feature map.

[0155] In order to further utilize the global context information of the support samples to optimize the features after interaction, the Squeeze-and-Co-excitation (SCE) mechanism is introduced to adaptively reweight the channel dimension of the initially enhanced features. Specifically, the masked support features Perform a global average pooling operation to obtain a channel descriptor vector Then, the vector is input into a small multi-layer perceptron (MLP), which consists of two fully connected layers and ReLu activation function, and outputs the weighted channel attention weight vector The weight calculation formula can be expressed as:

[0156]

[0157] in, is the final calculated channel attention weight vector, represents the global average pooling operation, Represents a multilayer perceptron network for learning nonlinear inter-channel dependencies from global channel descriptors and generating attention weights.

[0158] Finally, the calculated channel attention weights Support features applied to preliminary enhancements and query features , and then weighted by multiplying each channel to obtain the final enhanced support feature and enhanced query features The calculation formula is as follows:

[0159]

[0160]

[0161] Among them, among them, and They represent the final enhanced support features and query features respectively. As the index of all spatial locations, As the index of the feature channel, Indicates the The attention weight of each channel, Represents simple numeric multiplication.

[0162] (5) Prototype-loss compensation;

[0163] Specifically, after obtaining the enhanced support features, the model can use the features to generate the main support prototype to guide the segmentation of the query image. However, the pooling operation in the prototype generation process often leads to information loss. Therefore, the present invention uses an innovative prototype-loss compensation module to identify and compensate for the spatial detail information and difficult-to-distinguish sample information that may be lost in the generation process of the main prototype, thereby generating a compensated prototype and calculating an auxiliary loss at the same time to guide the model to better understand the support samples. Because the loss function of the model is usually calculated by summing up after each training round, the calculation of all losses (including auxiliary losses) in the present invention will be explained in detail later. The detailed processing flow of this module is as follows. Figure 5 shown.

[0164] First, the enhanced support features output by the bidirectional feature interaction module guided by the prior Generate a main support prototype. The main prototype is intended to capture the core and global feature representation of the support category. In a preferred embodiment of the present invention, the main support prototype is generated by enhancing the support features. This is done by performing a global average pooling operation. The calculation formula is as follows:

[0165]

[0166] in, Represents the calculated main support prototype.

[0167] Next, the module discovers information that the main prototype may have overlooked by performing a “self-prediction” on the support samples and generates a compensating prototype while calculating the auxiliary loss.

[0168] In order to make segmentation prediction on the support branch, it is necessary to construct appropriate input features. The input features are composed of the original support fusion features and the previously calculated main support prototype Specifically, the main support prototype By copying and expanding the spatial dimension, its spatial resolution is consistent with the original support fusion feature Then, the original support is integrated with the special With two copied and extended main prototypes The calculation formula for splicing in the channel dimension is:

[0169]

[0170] in, is the input feature used to support image prediction in this module, It means expanding the vector in the spatial dimension.

[0171] The features constructed above Input into the decoder to obtain the pixel-level segmentation prediction mask of the support image. The process can be described as:

[0172]

[0173] in is the pixel-level segmentation prediction mask of the support image under the guidance of the main prototype, The decoder uses the context-aware attention-guided feature aggregation decoding module proposed in this paper to predict the support image. This module shares parameters with the query branch's decoding module to avoid introducing additional parameters and reduce the risk of overfitting. This context-aware attention-guided feature aggregation decoding module will be described in detail in subsequent steps.

[0174] Based on the support image prediction mask and the true support foreground mask, all correctly predicted foreground pixels can be calculated using the following formula:

[0175]

[0176] in Gather all correctly predicted foreground pixels, representing the correctly segmented foreground area. ” represents the logical AND operation.

[0177] By removing the correctly segmented foreground area from the true support foreground mask, we can get all the incorrectly predicted foreground pixels, which can be expressed as:

[0178]

[0179] in, represents the foreground information lost during the prototype generation process, ” represents a remove operation.

[0180] In view of the spatial detail information and difficult-to-distinguish sample information that may be lost during the prototype generation process, this paper regards these lost foregrounds as a new defect sample. In order to provide more complete context information, this paper uses the surrounding background information of the supporting image to supplement the defect, so that the network can not only learn "what" is omitted, but also "where" it is omitted, thereby providing key information for improving the accuracy of subsequent predictions. The complete compensation mask is defined as follows:

[0181]

[0182] The compensation prototype can be obtained by mask average pooling using the compensation mask and fusion support features. The specific formula is as follows:

[0183]

[0184]

[0185] in, is After downsampling, Indicates that the bilinear interpolation downsampling operation will be Adjust to the characteristics The same spatial resolution, is the calculated compensation prototype.

[0186] (6) Context-aware attention-guided feature aggregation decoding;

[0187] Specifically, the present invention implements this process through a context-aware attention-guided feature aggregation decoding module. The core purpose is to effectively fuse the enhanced features from the query image itself with the two prototype features condensed from the support set information, and process them through this module to finally generate a pixel-level defect segmentation prediction result for the query image. Figure 6 Schematic diagram of the structure of the context-aware attention-guided feature aggregation decoding module in the present invention.

[0188] First, we need to enhance the query features , Main Support Prototype and compensation prototypes Fusion is performed to form the input of the module, since the prototype and The space size is Therefore, the prototype needs to be expanded in the spatial dimension so that its spatial resolution is consistent with the enhanced query feature. The fusion feature generated by this process is recorded as , and its calculation formula is:

[0189]

[0190] Subsequently, the Convolution operation generates compressed features ,in In order to capture multi-scale context information, we further The input is sent to the Atrous Spatial Pyramid Pooling (ASPP) module to extract context information. This module uses a set of dilated convolution kernels with different dilation rates to extract features at multiple scales. In addition to dilated convolution, a global average pooling operation is also used to aggregate the global context information into a vector. This vector is then passed through Convolution processing, and then expansion and fusion with multi-scale features. The context information extracted and spliced ​​by the ASPP module will also be compressed to form This process can be written as:

[0191]

[0192] in, Contains a global average pooling operation, a Convolutional layer and a dilation operation. Represents multiple parallel hole convolution layers, subscript is the void ratio, Represents the ReLU activation function.

[0193] To further enhance the feature representation, an attention-guided residual block with two branches is adopted: the convolution branch and the attention branch.

[0194] The convolution branch passes through two The convolutional layer processes the output of ASPP. This process can be expressed as:

[0195]

[0196] in represents the output of the convolution branch, express Convolutional layer.

[0197] At the same time, the attention branch adaptively weights features through the attention weights generated by the CoordAttention unit, thereby refining the features. Specifically, the CoordAttention unit captures long-range dependencies with precise positional information by decomposing the overall process into two parallel one-dimensional feature encoding processes along the width (x direction) and spatial height (y direction).

[0198] First, input features Average pooling will be performed in both the width and height directions, and the corresponding calculation formula is:

[0199]

[0200]

[0201] in, Indicates that the ASPP output feature map is at position The value at Represents the result of pooling along the width direction, representing the feature map in the height At , take the average of the features over all widths; Represents the result of pooling along the height direction, representing the feature map at height At , the features at all heights are averaged.

[0202] The two vectors are then concatenated and processed through a shared convolutional layer, batch normalization, and nonlinear activation function to encode precise location information. This process can be summarized as:

[0203]

[0204]

[0205] in, is a transpose operation, the purpose of which is to Adjust to the right shape so that it can be combined with Stitched together. It is the output feature after splicing and convolution processing. Represents the batch normalization operation, which is used to accelerate training and improve the generalization ability of the model; represents the nonlinear activation function, It is the output feature after batch normalization and nonlinear activation.

[0206] Next, the processed feature vector is split and fed into two independent convolutional layers and Sigmoid functions to generate attention weight maps along the height and width directions. The calculation formula for this process is:

[0207]

[0208]

[0209] in, Represents a split operation, They correspond to the information along the width and height directions after splitting. is the Sigmoid activation function, and Corresponding to the attention weights in the width and height directions respectively.

[0210] The attention weight output by the coordinate attention unit is expressed as follows:

[0211]

[0212] Then the output of the attention-guided residual block of the two branches is expressed as:

[0213]

[0214] The final predicted segmentation mask for the query image can be obtained through a fully convolutional decoder:

[0215]

[0216] in, is the final prediction result of the query image, Represents a fully convolutional decoder, consisting of a Convolutional layer, a ReLU function activation layer, a Convolutional layer. represents a bilinear interpolation upsampling operation used to adjust the predicted mask to the same spatial resolution as the query image, Used to obtain the segmentation mask, which means using this function to determine the category to which each pixel belongs from the output of the model, thereby generating the final segmentation mask.

[0217] (7) Loss function calculation;

[0218] After a training round, the loss function between the predicted segmentation result and the true label is calculated, and back propagation is performed to optimize the model parameters; specifically, the predicted segmentation results mentioned in the present invention are divided into two types: the segmentation prediction result of the support image and the segmentation prediction result of the query image. The segmentation prediction result of the support image is output by the prototype-loss compensation module, and the loss calculated with the true support image label is used as the auxiliary loss. The segmentation prediction result of the query image is output by the context-aware attention-guided feature aggregation decoding module, and the loss calculated with the true query image label is used as the main loss. The present invention uses the cross entropy loss function as the loss calculation function to calculate the model loss. The calculation formulas for the main loss, auxiliary loss, and total loss are as follows:

[0219]

[0220]

[0221]

[0222] in, and denote the cross entropy loss in query image segmentation prediction and support image segmentation prediction, respectively. is the ratio of auxiliary loss, in the present invention . and Represent the segmentation prediction probability maps corresponding to the support image and the query image respectively. Note that and ,in The function finds the index of the class with the highest probability at each pixel location. represents the number of pixels in the probability map, is the index of the pixel, represents the number of samples in the support set, is the index of the sample.

[0223] Step S4: After the training is completed, the support set and the query image to be tested are input into the converged segmentation network, and the query image prediction result is output;

[0224] Specifically, during the training and testing phases, image processing and segmentation procedures are identical. The difference is that during the training phase, the model parameters are optimized under the supervision of the query image's true labels. During the testing phase, the model receives images and their categories that have never been seen during training. To test the model's performance and generalization, the query image's true labels are only used to evaluate the accuracy of the model's predictions. In short, the model parameters are not optimized during the testing phase and are only used to make predictions.

[0225] The following comparative experiments will demonstrate the beneficial effects of the present invention.

[0226] The dataset used in this comparative experiment is the FSSD-12 dataset. Specifically, the FSSD-12 dataset contains images of 12 types of metal surface defects, including wear, iron dust, liquid stains, plate scale, oil stains, water stains, repairs, holes, red iron, roller marks, scratches, and inclusions. To avoid long-tail distribution issues, each category contains 50 images and corresponding pixel-level annotation masks, and each defect image is classified into only one category. All images are uniformly resized to 200×200 pixels.

[0227] This comparative experiment uses the most commonly used metric in small-shot segmentation: the mean intersection over union (MIoU). The intersection over union (IoU) is calculated as the intersection of the predicted foreground area and the ground-truth foreground area divided by the union of their two areas. The mean IoU is calculated by calculating the IoU for each class, then summing the IoU values ​​for all classes and dividing the sum by the total number of classes, C.

[0228] The calculation formulas for the intersection-and-union ratio and the average intersection-and-union ratio are as follows:

[0229]

[0230]

[0231] in, represents the total number of categories, Indicates the The intersection-over-union ratio of the categories. (True Positive) refers to the In a category, the model predicts that the number of foreground pixels belongs to this category and the true label also belongs to this category. (False Positive) refers to the The number of background pixels in a category that the model predicts as belonging to the category of foreground pixels, but the true label is the category of background pixels. (False Negative) refers to the The number of foreground pixels that the model predicts to belong to the background class, but the true label belongs to that class. The value range is 0 to 1, where 0 means that the foreground area of ​​the predicted result does not overlap with the foreground area of ​​the true label, and 1 means that the foreground area of ​​the predicted result completely overlaps with the foreground area of ​​the true label. It can comprehensively evaluate the accuracy and completeness of the segmentation results. The higher the value, the better the segmentation performance.

[0232] The following comparisons show the proposed method with current mainstream methods for small-sample surface defect segmentation and small-sample semantic segmentation, including TGRNet (Method 1), BAM (Method 2), CPANet (Method 3), and PFENet++ (Method 4). All compared methods use ResNet50 as the backbone network. Experiments are conducted in both 1-shot (support set contains one annotated image) and 5-shot (support set contains five annotated images) settings to comprehensively evaluate the performance of each method.

[0233] The TGRNet method comes from the literature: Bao Y, Song K, Liu J, et al. Triplet-graphreasoning network for few-shot metal generic surface defect segmentation[J]. IEEE Transactions on Instrumentation and Measurement, 2021, 70: 1-11.

[0234] The BAM method comes from the literature: Lang C, Cheng G, Tu B, et al. Learning what not tosegment: A new perspective on few-shot segmentation[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022: 8057-8067.

[0235] The CAPNet method comes from the literature: Feng H, Song K, Cui W, et al. Cross position aggregation network for few-shot strip steel surface defect segmentation[J]. IEEE Transactions on Instrumentation and Measurement, 2023, 72: 1-10.

[0236] The PFENet++ method comes from the literature: Luo X, Tian Z, Zhang T, et al. PFENet++:Boosting Few-Shot Semantic Segmentation With the Noise-Filtered Context-AwarePrior Mask[J]. IEEE Transactions on Pattern Analysis and MachineIntelligence, 2024, 46(2): 1273-1289. Table 1 Comparative test results

[0237] The comparative experimental results in Table 1 clearly show that our method achieves significant performance improvement on the FSSD-12 dataset. Compared with the other four comparison methods, our method achieves the highest MIoU score in both 1-shot and 5-shot settings, demonstrating its significant advantage in small-sample defect segmentation tasks.

[0238] An embodiment of the present invention also provides a computer program product, comprising computer program instructions, which are stored on a computer-readable storage medium. When the instructions are executed by a processor, a method for segmenting surface defects of an object based on a small sample as described in the above technical solution is implemented.

[0239] An embodiment of the present invention also provides an electronic device, comprising a memory and a processor communicatively connected to the memory, wherein the memory stores computer program instructions, and when the processor executes the computer program instructions, it implements a method for segmenting surface defects of an object based on a small sample as described in the above technical solution.

[0240] Finally, it should be noted that the above embodiments are intended to illustrate the technical principles of the present invention and should not be construed as limiting the scope of protection of the present invention. Guided by the design principles of the present invention, those skilled in the art may make various improvements and variations to the above embodiments. Such improvements and variations should fall within the scope of protection claimed by the present invention, which is clearly defined by the claims.

Claims

1. A method for segmenting surface defects of an object based on a small sample, characterized in that: The following steps are involved: Step S1: collect images of surface defects of different categories of objects and their corresponding segmentation masks, construct a small sample data set and divide it into a training set and a test set; Step S2: construct training and test tasks, and extract corresponding support sets and query sets for each task; Step S3, constructing and training a defect segmentation network, wherein the defect segmentation network includes a query prior mask generator, a prior-guided bidirectional feature interaction module, a prototype-loss compensation module, and a context-aware attention-guided feature aggregation decoding module; The query priori mask generation module is used to generate a query image priori mask map; The prior-guided bidirectional feature interaction module is used to obtain enhanced support features and enhanced query features, and generate support prototypes; The prototype-loss compensation module predicts the support image using the support prototype to obtain the predicted segmentation result of the support image and generates a compensation prototype; The context-aware attention-guided feature aggregation decoding module is used to fuse the enhanced query features with the supporting prototype and the compensating prototype to ultimately generate a pixel-level defect segmentation prediction result for the query image; Step S4: After the training is completed, the support set and the query image to be tested are input into the converged defect segmentation network, and the query image prediction result is output.

2. The method for segmenting surface defects of an object based on a small sample according to claim 1, characterized in that: The process of defect segmentation network is as follows: given a pair of support samples and query image , using two weight-sharing feature extraction networks to extract multi-level support features and query features ,in, and They are the support image in the support set and its corresponding segmentation mask, and They represent the feature extractor The feature map output by the layer; for subsequent processing, the third and fourth level features are spliced ​​and fused to generate , ; Using the fourth level features to generate the query image prior mask and the query prior mask generator and After filtering respectively, the prior-guided bidirectional feature interaction module is used to ensure sufficient feature interaction between the two features and generate enhanced support features respectively. and enhanced query features ; Then, through global average pooling, using Generate support prototype At this time, the prototype-loss compensation module uses the support prototype to predict the support image and obtain the predicted segmentation result of the support image , thereby detecting the spatial information lost during the generation of the main prototype and generating a compensation prototype , and provide auxiliary loss; finally, the context-aware attention-guided feature aggregation decoding module is integrated with and the prototype from the support branch Finally, the features are aggregated and decoded to predict the query image , generate segmentation results .

3. The method for segmenting surface defects of an object based on a small sample according to claim 2, characterized in that: Using the feature extraction network that has been pre-trained on the ImageNet dataset and whose parameters are frozen, the support image and query image in the current training or test task are processed respectively, and the third-level features and the fourth-level features are preferably extracted for subsequent feature processing; wherein, the third-level support and query features are respectively denoted as and , the support and query features of the fourth level are respectively recorded as and ; The generation process of fusion features is as follows: ; ; in, and Represent the calculated support fusion features and query fusion features respectively, Indicates concatenation of the two input feature maps in the channel dimension. Represents a The convolution layer receives the channel-joined feature map as input and adjusts its channel number to the preset fusion channel number through convolution operation.

4. The method for segmenting surface defects of an object based on a small sample according to claim 2, characterized in that: The specific implementation process of querying the priori mask generator in step S3 includes: Generate foreground prototype and background support prototype: ; ; in, and Indicates the fourth level of support and query features, represents the masked average pooling operation, is the spatial resolution of the fourth-level features, and are the calculated foreground and background prototype vectors, and They are respectively Downsampled foreground and background support masks with consistent spatial resolution; and ,in Indicates that the original image mask is adjusted to the same spatial resolution as the fourth-level feature through bilinear interpolation downsampling operation. To support mask; is the index of all spatial locations, represents element-wise multiplication; Measuring Level 4 Support Characteristics The similarity between the features of each spatial position in and the support prototype calculated above: ; ; in, and Represent the similarity graphs of the support vector to the foreground and background of the query image, respectively. is the index of all spatial positions, symbol Represents a vector norm, is a very small positive number added to the denominator to ensure numerical stability, and Respectively represent taking the minimum value and taking the maximum value; Normalize the two original similarity graphs: ; ; in, and denote the initial foreground and background prior masks of the query image, respectively; Perform the blur region identification and removal steps. By calculating the Hadamard product of the two normalized similarity maps, we can identify the regions composed of blurred pixels. These regions are then removed from the foreground prior mask to obtain the purified prior mask. Finally, the query prior mask is normalized again to obtain the final required query prior mask: ; ; ; in, Indicates the fuzzy area, Represents the purified mask, where "-" represents the removal operation. Indicates that the bilinear interpolation upsampling operation will Adjust to the same spatial resolution as the query image to obtain the final query prior mask.

5. The method for segmenting surface defects of an object based on a small sample according to claim 2, characterized in that: Specific implementation process of the prior-guided bidirectional feature interaction module in step S3 include: Perform feature masking operations to reduce the interference of complex background. The two feature masking operations can be expressed as: ; ; in, Represents the supported mask The support features after filtering, Represents the query prior mask The filtered query features, and They are respectively and Downsampling mask with consistent spatial resolution; The masked support features and the masked query features obtained after preprocessing are input into the bidirectional cross attention unit to establish pixel-level correspondence and achieve mutual feature enhancement. The process uses the following formula to enhance the query and support features: ; ; in, and Represent the support features and query features that are initially enhanced after interaction, 、 、 Both represent convolution, represents matrix multiplication, are the indices of all possible positions, Represents a feature map After convolution mapping at position The eigenvectors on Represents another feature map passing through After convolution mapping at position The eigenvectors on , Indicates that the similarity between two feature vectors is calculated by dot product operation. is the number of positions in the input feature map that participate in the attention calculation; In order to further utilize the global context information of the supporting samples to optimize the features after interaction, a squeezing and co-excitation mechanism is introduced to adaptively reweight the channel dimension of the initially enhanced features. The weight calculation formula can be expressed as: ; in, is the final calculated channel attention weight vector, represents the global average pooling operation, represents a multi-layer perceptron network used to learn nonlinear inter-channel dependencies from global channel descriptors and generate attention weights; Finally, the calculated channel attention weights Support features applied to preliminary enhancements and query features , and then weighted by multiplying each channel to obtain the final enhanced support feature and enhanced query features ; The calculation formula is as follows: ; ; Among them, among them, and Represent the final enhanced support features and query features, As the index of all spatial locations, As the index of the feature channel, Indicates the The attention weight of each channel, Represents simple numeric multiplication.

6. The method for segmenting surface defects of an object based on a small sample according to claim 2, characterized in that: The specific implementation process of the prototype-loss compensation module generating the compensation prototype in step S3 includes: The enhanced support features output by the prior-guided bidirectional feature interaction module generate the main support prototype, and its calculation formula is as follows: ; in, Represents the calculated main support prototype; The main support prototype By copying and expanding the spatial dimension, its spatial resolution is consistent with the original support fusion feature Then, the original support is integrated with the special With two copied and extended main prototypes The calculation formula for splicing in the channel dimension is: ; in, are input features used to support image prediction, Indicates that the vector is expanded in the spatial dimension; Perform a self-prediction on the support samples to discover information that the main prototype may have overlooked and generate a compensating prototype: ; in is the pixel-level segmentation prediction mask of the support image under the guidance of the main prototype, The decoder is a context-aware attention-guided feature aggregation decoding module; Based on the support image prediction mask and the true support foreground mask, all correctly predicted foreground pixels are calculated using the following formula: ; in Gathers all correctly predicted foreground pixels, representing the correctly segmented foreground area, " represents the logical AND operation; Remove the correctly segmented foreground regions from the true support foreground mask to get all incorrectly predicted foreground pixels: ; in, represents the foreground information lost during prototype generation," " represents the removal operation; The surrounding background information of the supporting image is used to supplement the defect. The complete compensation mask is defined as follows: ; The compensation prototype is obtained by mask average pooling using the compensation mask and fusion support features. The specific formula is as follows: ; in, is After downsampling, Indicates that the bilinear interpolation downsampling operation will be Adjust to the characteristics The same spatial resolution, is the calculated compensation prototype.

7. The method for segmenting surface defects of an object based on a small sample according to claim 2, characterized in that: The implementation process of step S3 specifically includes: Enhanced query features , Main Support Prototype and compensation prototypes To perform fusion, take as input: ; use Convolution operation generates compressed features ,in ; In order to capture multi-scale context information, we further Input into the void space pooling pyramid module to form , this process is written as: ; in, Contains a global average pooling operation, a Convolutional layer and a dilation operation; Represents multiple parallel hole convolution layers, subscript is the void ratio, Represents the ReLU activation function; In order to further enhance the feature representation, an attention-guided residual block with two branches is adopted, namely the convolution branch and the attention branch. The output of the attention-guided residual block of the two branches is expressed as: ; The final predicted segmentation mask for the query image is obtained through a fully convolutional decoder: ; in, is the final prediction result of the query image, Represents a fully convolutional decoder, consisting of a Convolutional layer, a ReLU function activation layer, a Convolutional layers, represents a bilinear interpolation upsampling operation used to adjust the predicted mask to the same spatial resolution as the query image, Used to obtain the segmentation mask, which means using this function to determine the category to which each pixel belongs from the output of the model, thereby generating the final segmentation mask.

8. The method for segmenting surface defects of an object based on a small sample according to claim 2, characterized in that: The loss function calculation formula used in training the defect segmentation network in step S3 is specifically: ; ; ; in, and denote the cross entropy loss in query image segmentation prediction and support image segmentation prediction, respectively. From Prototype-Loss Compensation Module, Auxiliary loss The proportion of and Represent the segmentation prediction probability maps corresponding to the support image and the query image respectively, and ,in The function finds the category index with the highest probability at each pixel position; represents the number of pixels in the probability map, is the index of the pixel, represents the number of samples in the support set, is the index of the sample.

9. A computer program product, characterized in that The method comprises computer program instructions, which are stored on a computer-readable storage medium. When the instructions are executed by a processor, the method for segmenting surface defects of an object based on a small sample is implemented as described in any one of claims 1 to 8.

10. An electronic device, characterized in that: It includes a memory and a processor in communication with the memory, wherein the memory stores computer program instructions, and when the processor executes the computer program instructions, it implements a surface defect segmentation method based on a small sample according to any one of claims 1 to 8.

Citation Information

Cited By

  • Pipeline crack detection method, device, equipment and medium under condition of few sample data

    CN121388406A

  • A method, apparatus, equipment and medium for detecting pipeline cracks using limited sample data.

    CN121388406B