Small sample semantic segmentation method based on visual large model multi-scale prompt
Through the multi-scale prompt method of visual large models, the problem of insufficient generalization ability of deep learning models for new categories under small sample conditions is solved, and efficient semantic segmentation of small samples is achieved, and the generated target prompt coding is more accurate and interpretable.
Patent Information
- Application Number
- CN202510113896.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Existing deep learning models lack the ability to generalize new categories under small sample conditions, especially in semantic segmentation tasks, requiring a large amount of labeled data and being costly.
Using a multi-scale prompt method based on visual big model, a similar support query sample pair is constructed, and a multi-scale feature extraction module and a multi-scale prompt generation module are used to generate multi-scale prompt encoding, and combined with visual big model SAM to perform semantic segmentation of small samples.
It improves the generalization ability of new classes that have not been seen during the training process, improves the accuracy and efficiency of semantic segmentation of small samples, and the generated target hint encoding has better interpretability and accuracy.
Smart Images

Figure CN120107579A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision semantic segmentation, and in particular relates to a small sample semantic segmentation method based on multi-scale prompts of a large visual model. Background Art
[0002] Neural network models based on deep learning have achieved remarkable results in many fields, but their success is based on a large amount of labeled data. For those categories that are rarely or not included in the training data, the effect of deep neural networks is often greatly reduced, and the generalization ability for new categories is often not very strong. For downstream tasks, it is usually necessary to collect data again and use the collected relevant data to retrain or fine-tune the original model. However, it is very costly to collect a large amount of labeled data, especially for intensive prediction tasks such as semantic segmentation. The cost is even higher, and it also takes a lot of time. Over-reliance on data greatly limits the application of deep learning models. In order to alleviate the dependence of deep learning networks on data, small sample learning aims to train a network that can generalize well to new categories with scarce labeled training data, and can complete the specified task given only a small number of labeled samples. In other words, it is urgent to train a network with learning ability, which can learn a new concept with a small amount of learning materials, so as to alleviate the problem that deep learning networks are too dependent on training data.
[0003] The main method of the small sample semantic segmentation model is to complete the matching of pixel position features through prior knowledge. The current small sample semantic segmentation method is mainly based on the pre-trained backbone network to obtain universal capabilities for all categories. Such prior knowledge is not necessarily optimal for downstream segmentation tasks. Therefore, by introducing a large visual model for segmentation tasks, the segmentation accuracy can be further improved.
[0004] There are two main methods to use large visual models to improve the accuracy of small sample semantic segmentation. One is to use the preliminary segmentation results of the target image as prior knowledge and introduce them into the model for use; the other is to generate visual cues for the current target and input them into the large visual model to complete the segmentation task. Summary of the invention
[0005] The main purpose of the present invention is to utilize the general capabilities of a large visual model to improve the segmentation accuracy of a target object and to provide a small sample semantic segmentation method based on multi-scale cues from a large visual model.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions: A small sample semantic segmentation method based on multi-scale prompts of a large visual model, the small sample semantic segmentation method comprising the following steps: S1. Use the public semantic segmentation dataset to build similar support query sample pairs. The similar support query sample pairs include query images. , Support images and support mask ; S2, through the fusion feature extraction module The support fusion features and the query fusion features are obtained respectively, and the visual prior of the query image is obtained at the same time, wherein the fusion feature extraction module It includes a backbone network, a pair of parallel feature fusion convolutional layers, and a pair of parallel feature activation convolutional layers connected in sequence; S3, multi-scale feature extraction module Model and downsample the support fusion features and query fusion features respectively, and extract multi-scale support features and query features respectively; S4. Multi-scale prompt generation module The multi-scale support features and query features are used to generate multi-scale prompt codes, and the multi-scale prompt codes are fused from bottom to top to generate target prompt codes; wherein, the process of step S4 is as follows: S41, inputting query features and supporting features of the current scale, obtaining similarity graphs of the query features and supporting features using cosine similarity, adjusting the similarity graph using intermediate results of the next scale, aggregating the input learnable prompt vector using the adjusted similarity graph, and generating similarity prompt encoding of the current scale; S42, fusing the similar prompt code of the current scale with the prompt code and the intermediate result of the next scale to generate the prompt code of the current scale, and outputting the prompt code and the intermediate result generated by it to the previous scale to continue iteration until the target prompt code is generated; S5. Input the target hint code and the query image into the hint-type visual large model SAM to obtain the segmentation result of the query image, wherein the visual large model SAM includes an encoder and a decoder, input the query image into the encoder to obtain the query image code, and input the target hint code and the query image code into the decoder to obtain the segmentation result; construct a prediction model, which includes a fusion feature extraction module connected in sequence , Multi-scale feature extraction module , Multi-scale prompt generation module The prediction model is trained using the intermediate results and the final results to form a total loss function. During the training process, the parameters of the visual large model SAM remain frozen. S6. Input the supporting query sample pairs in the small sample data into the trained prediction model, segment the query image, and calculate the segmentation accuracy based on the prediction results.
[0007] Furthermore, the process of constructing similar supporting query image sample pairs in step S1 is as follows: The public datasets of semantic segmentation are divided into four groups by category, of which three training sets are used for training and the remaining one test set is used for testing. In each set of data, the training set and the test set contain a support set and a query set, respectively. The support set consists of support images and their corresponding densely annotated masks, and the query set consists of single images of the same type and their masks. The densely annotated masks of the support images in the support set are used as prior knowledge. During the training process, the prediction network is supervised by the densely annotated masks of the query set images to optimize the network parameters. In the testing phase, the densely annotated masks of the query images are used as a standard to evaluate the model performance.
[0008] Furthermore, the process of obtaining the supporting fusion features and the query fusion features and the query visual prior in step S2 is as follows: S21. Use the pre-trained ResNet-50 as the backbone network. ResNet is a deep residual network, which is a deep learning algorithm that uses cross-layer connections to obtain residuals. ResNet-50 is a specific implementation of the deep residual network and is widely used in image classification and segmentation. For details, please refer to the paper KM He, XY Zhang, SQ Ren,SQ, and J. Sun, Deep Residual Learning for Image Recognition[C], in Proc. IEEEConf. Comput. Vis. Pattern Recognit., Jun. 2016, pp. 770-778. Select the outputs of the 2nd and 3rd layers in ResNet-50 as the intermediate layer features, and the outputs of the 4th and 5th layers as the high-level features. Because ResNet-50 is pre-trained based on classification tasks, the size of the intermediate layer features is larger and contains more semantic information, while the size of the high-level features is smaller, which is beneficial for classification tasks. For intensive segmentation tasks, directly using high-level features cannot improve performance. The query images and supporting images Enter the backbone network and obtain features:
[0009]
[0010] in, represents the pre-trained backbone network, , Respectively represent The query features and support features output by the layer, Indicates The number of channels of layer features, Indicates The height of the layer features, Indicates The width of the layer feature; Query the middle layer features respectively and support for mid-level features Input to each feature fusion convolution layer for dimension reduction and feature fusion to obtain query intermediate features and support intermediate features ; Using high-level features and Generate visual priors by first using bilinear interpolation to support the mask The size is adjusted to support the intermediate features Consistent, using the Hadamard product method to remove high-level support features Irrelevant background features:
[0011] in, To support binary masking of images, “⊙” represents Hadamard product; Bilinear interpolation is a mathematical method that performs linear interpolation by considering the four pixel values closest to the interpolation point and assigning different weights to each point according to its distance from the interpolation point; The correlation between pixels is discovered by calculating the cosine similarity of the query feature and the supporting feature at each pixel position. :
[0012] in" " represents the 3D vector inner product operation at all feature positions. represents the L2 norm; For the query feature of each position, take its maximum similarity value as the prior value of the current position, and use the minimum-maximum normalization method to normalize its value to between 0 and 1 to obtain the visual prior ; The visual prior is essentially a rough segmentation result of the target object in the query image. Using the visual prior can make each module pay more attention to the area where the target object may exist, thus improving the segmentation accuracy. S22, using bilinear interpolation method to support mask The size is adjusted to support the intermediate features Consistently, the prototype vector of the current category is obtained using the mask average pooling operation:
[0013] in, is the prototype vector of the current category, Indicates support for intermediate features In Location The characteristic value of To support the mask In Location The annotation value of is a conditional function that takes 1 if the value in the brackets is true, and 0 otherwise; The intermediate features will be queried , prototype vector and visual priors Splicing along the channel dimension, and then passing through the query feature activation convolution layer to obtain the query fusion feature ; Will support intermediate features , prototype vector Splicing along the channel dimension, and then passing through the support feature activation convolution layer to obtain the support fusion feature .
[0014] Furthermore, the process of extracting multi-scale features by the multi-scale feature extraction module in step S3 using fusion features is as follows: The multi-scale feature extraction module includes Each layer of attention blocks is connected in sequence, and the output of each layer of attention blocks is the feature information at a single scale. The outputs of the layer attention blocks are combined to form a multi-scale feature pyramid; each attention block includes a local attention block LocalAttnBlock and a self-attention block SelfAttnBlock connected in sequence. The formal representation of a single layer is as follows:
[0015] in, Indicates the current layer number. It is The features of the layer output, , It is The features of the layer output, , and there is , , , Indicates The number of channels of the layer pyramid feature, Indicates The layer pyramid features high, Indicates The width of the layer pyramid feature; The local attention block is good at local modeling, but weak at global modeling; the self-attention block is good at global modeling, but weak at local modeling; by combining the local attention block with the self-attention block, the advantages can be complemented, so that the multi-scale feature extraction module performs well in both local and global modeling; Each local attention block is essentially a multi-branch convolution layer, which obtains feature information under different receptive fields through convolution kernels of different sizes; the basic operation of each self-attention block is the attention mechanism :
[0016] in, is the query vector, is the key vector, is a value vector, To query features, is the key feature, is the value feature, is the corresponding projection matrix, is the value of the feature dimension; if , , If they are the same feature, it is called self-attention, if they are different, it is called mutual attention; The operation is applied to each row of the matrix to convert the attention scores into a probability distribution. , is the input matrix, For the matrix In Location The value of The query fusion features and support fusion features are input into the multi-scale feature extraction module respectively to obtain the multi-scale query feature pyramid and multi-scale support feature pyramid ,in, Represents the query feature pyramid The characteristics of the layer, Indicates support feature pyramid Characteristics of the layer; The feature sizes of each layer in the feature pyramid are different, so the features at each scale preserve different semantic information. Specifically, the size of low-level features is larger and can retain more local detail information, while the size of high-level features is smaller and contains more abstract semantic information. Using the rich semantic information at multiple scales can construct more accurate prompt information.
[0017] Furthermore, the multi-scale prompt generation module in step S4 Include The output of each layer is the hint code at the current scale. Each mutual attention hint code generation block first mines the similar information between the features of the current scale according to the cosine similarity, and then aggregates the hint code and intermediate results of the next scale to generate the hint code at the current scale. The mutual attention cue coding generation blocks at different layers input feature information at different scales, so the generated similarity cue coding focuses on different things; specifically, the similarity cue coding generated using larger features pays more attention to local details, targeting more distinguishable areas in the target object rather than the overall area, and the granularity of the cue is smaller; the similarity cue coding generated using smaller features pays more attention to overall abstract information, generates rougher cue information, but may lack some details, and the granularity of the cue is larger; using a bottom-up optimization approach can make the focus of the generated cue coding gradually transition from the overall to the local, and from the abstract to the details.
[0018] Furthermore, the process of generating the similarity hint code of the current scale in step S41 is as follows: For Layer characteristics and , first reshape the features and shape, Reshape into , and obtain the reshaped features respectively and Because the hint encoding should only be related to the similarity relationship between the query feature and the supporting features, but not directly related to a single feature, and the data distribution of the hint encoding is relatively independent of the data distribution of the query feature and the supporting feature, a set of learnable hint vectors is introduced. , to decouple the data distribution between the prompt encoding and the features, and similarly convert the prompt vector Reshape , respectively calculate the query vector in the attention mechanism , key vector , value vector , get the query vector , key vector Sum value vector , where the query vector From the query features , the key vector From the supporting features , value vector A learnable hint vector from a set of inputs , formally expressed as follows:
[0019] in, , , They are , , The projection matrix; In order to obtain the similarity relationship between each pixel position, calculate and The cosine similarity between them gives a similarity graph matrix :
[0020] in," " represents the vector inner product operation, represents the L2 norm of the vector, is the temperature hyperparameter, which can be ; Along The first dimension Operation, get the normalized similarity graph matrix :
[0021] in, For the matrix In Location The value of For the matrix In Location The value of Masks will be supported using bilinear interpolation Adjust to the corresponding size and reshape to get ,use Remove the similarity values of irrelevant background positions to obtain the similarity matrix after masking the background. ; Under multi-scale conditions, the mask score output by the next layer is used as a similar prior to adjust the attention map and gradually optimize the segmentation result. The formal expression is as follows:
[0022] in, For the Layer similarity prior, obtained by resizing the logits score of the segmentation mask output by the next layer through bilinear interpolation. A logits score greater than 0 indicates foreground, and less than 0 indicates background. The larger the absolute value, the greater the probability. It is a combination of a matrix flattening operation and a dimension expansion copy operation, which first compresses the vector into a one-dimensional shape. , and then expand and replicate along the second dimension until its shape is the same as Consistency; Using the adjusted similarity matrix Aggregate the input learnable hint vector to get the first Similarity cue encoding at the layer ; By actively introducing the similarity prior of the next scale to adjust the similarity matrix, the generation of similarity hint encoding is integrated into the key information of the next scale, paying more attention to the key areas that may contain the target object; the addition operation reflects the similarity matrix and similar priors The equal relationship between the two is similar to a "voting" mechanism, which judges the final similarity by combining the results of the similarity matrix and the similarity prior, thereby generating a more accurate similarity prompt encoding.
[0023] Furthermore, the process of generating the prompt code of the current scale in step S42 and continuing to iterate to generate the target prompt code is as follows: S421. If the current layer is the last layer, that is, , then directly use the similar prompt encoding at the current scale As a hint encoding at the current scale ,Right now = ; This prompt code After bilinear interpolation and resizing, it is directly input into the decoder of the visual large model SAM to obtain the intermediate result score of the last layer of segmentation , As a similarity prior input to the In the layer mutual attention hint encoding generation block, the hint encoding Also enter the Iterative optimization in layers; S422. If the current layer is not the last layer, , we need to aggregate the next layer of hint codes and optimize them step by step from bottom to top. Hint encoding of layers Use bilinear interpolation to upsample to a shape consistent with the scale of the current layer. Layer Similarity Hint Encoding Spliced together, input into a fusion layer for information fusion, and obtain the prompt encoding at the current scale ; After resizing through bilinear interpolation, it is input into the decoder of the visual large model SAM to obtain the current The intermediate result score of the layer segmentation , enter it into The mutual attention hint encoding is used as a similarity prior in the generation block, hint encoding Also enter the Iterative optimization in layers; The fusion layer is essentially a multi-layer perceptron MLP, which is two linear layers connected in sequence:
[0024] in, is a nonlinear activation function, which can be used Activation function, , and is a constant; and is the weight parameter of MLP; Repeat step S422, after After the iteration of the layer, the prompt encoding of the first layer output It not only aggregates the prompt codes at multiple scales, but also is obtained through continuous iterative optimization of the intermediate results of subsequent layers, making full use of the semantic information of different granularities at multiple scales as the output target prompt code.
[0025] Furthermore, in step S5, the query image and the final hint code are input into the hint-based visual large model SAM to obtain the segmentation result, and the process of constructing and training the prediction model is as follows: S51, input the query image into the encoder of the visual large model SAM to obtain the query image code, and then output the segmentation result of the query image according to the hint code and the query image code; input the multi-scale hint code generated in step S4 together with the query image code into the decoder of the visual large model SAM to generate a multi-scale segmentation result ; Select As the final prediction output , the rest of the intermediate results Used as similarity priors to feed into the multi-scale cue generation module middle; SAM is a "segmentation everything model" that can provide general image segmentation capabilities. Its core goal is to be able to perform object segmentation in any type of image without the need for specific training data sets or task adjustments. SAM mainly consists of two parts: an image encoder and a decoder. It uses a new hint method to accurately segment the target object in the image, but SAM itself does not have specific semantic information and only has general segmentation capabilities. Therefore, for downstream tasks such as small sample semantic segmentation, specific hint information needs to be provided. SAM is widely used in many segmentation or target detection tasks. For details, please refer to the paper A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, AC Berg, W.-Y. Lo, et al., "Segment anything," in Proceedings of the IEEE / CVF International Conference on Computer Vision, 2023, pp. 4015–4026. The present invention uses a specific implementation of SAM, SAMv2, as the basic model. S52, successive fusion feature extraction modules , Multi-scale feature extraction module , Multi-scale prompt generation module And the visual big model SAM builds a prediction model; according to the final prediction results and intermediate results Constructing the total loss function , with the loss function The prediction model constructed for optimization goal training, loss function The calculation method is as follows:
[0026] in, is the total cross entropy loss, which can evaluate the classification accuracy of the foreground and background of the output results and is used to continuously optimize the intermediate results and the final prediction results. , For the mask of the query image, the cross entropy loss function , is the true label, is the predicted probability; The distillation loss of the similarity matrix between the layers of the multi-scale hint generation module is used to make the similarity relationship between high-level abstract features as close as possible to the similarity relationship between low-level details. , KL divergence loss function , is the true distribution right The probability of is the estimated distribution right The probability of For the The similarity matrix of the layer is masked average pooled and normalized along the first dimension to obtain the similarity vector; The DICE loss for the final prediction result is related to a generalized definition of intersection-over-union, which can more directly optimize the segmentation result. , is the pixel position The probability value predicted as the result can be taken as , , is the true value of the pixel position label; , , are the weight hyperparameters of each loss respectively; Through the total cross entropy loss , distillation loss and DICE loss The output target hint encoding can make full use of semantic information of different granularities at multiple scales, and the hint information is more accurate.
[0027] Furthermore, the process in step S6 is as follows: S61. Based on the trained prediction model, the segmentation result of the query image is obtained through the forward propagation of the network; if the final segmentation mask logits score is greater than 0, it is considered to be the foreground, otherwise it is predicted to be the background; S62, calculating the segmentation accuracy according to the prediction results, and using the average intersection over union (mIoU) as an evaluation indicator of the segmentation accuracy; mIoU is a widely used evaluation metric for semantic segmentation tasks, especially in multi-category segmentation, because it balances the performance of each category regardless of the category frequency; mIoU is the average of the intersection over union (IoU) of all categories, which can comprehensively reflect the segmentation performance of the model on known categories; the calculation formula is as follows:
[0028]
[0029] in, For Category The number of true positive examples, i.e., correctly segmented pixels; For Category The number of false positives, i.e. pixels misclassified as class 𝑖; For Category The number of false negatives is the number of pixels that are actually of class 𝑖 but are not correctly classified; is the total number of categories.
[0030] Compared with the prior art, the present invention has the following advantages and beneficial effects: (1) The present invention proposes a method of constructing a multi-scale feature extraction module by combining local attention blocks with self-attention blocks, which achieves complementary advantages and enables the multi-scale feature extraction module to perform well in both local modeling and global modeling. The generated multi-scale feature pyramid can retain richer and more accurate semantic information.
[0031] (2) The multi-scale cue generation module designed in the present invention is a bottom-up iterative network structure that can use the intermediate prediction results of the next layer as similarity priors to optimize the results of the current layer, and can fuse the cue encoding from coarse-grained to fine-grained, gradually transitioning the focus of the cue encoding from the overall to the local, and from the abstract to the details, thereby improving the accuracy of the target cue encoding.
[0032] (3) The method proposed in the present invention combines the target hint coding with the large visual model, making full use of the prior knowledge and general capabilities of the large visual model, improving the generalization ability of new classes that have not been seen in the training process, and can effectively alleviate the problem of generalization performance degradation of deep learning; compared with other methods for generating hints, the target hint coding generated by the present invention has better interpretability, and the information at each position of the target hint coding can correspond one-to-one with the prediction result, which is more convenient for optimizing the result.
[0033] (4) The method based on multi-scale prompts of a large visual model proposed in the present invention can achieve better performance with fewer iterative training times in the task of small-sample semantic segmentation by continuously utilizing intermediate results for iterative optimization, thereby improving the accuracy of small-sample semantic segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0035] Figure 1 It is a flowchart of a small sample semantic segmentation method based on multi-scale prompts of a large visual model disclosed in the present invention; Figure 2 A schematic diagram of the structure of a multi-scale feature extraction module and a multi-scale prompt generation module and the connection structure therebetween in an embodiment of the present invention; Figure 3 This is a training flow chart in an embodiment of the present invention. DETAILED DESCRIPTION
[0036] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.
[0037] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0038] Example 1 Figure 1 is a flowchart of a small sample semantic segmentation method based on multi-scale prompts of a large visual model disclosed in an embodiment of the present invention, such as Figure 1 As shown, the method comprises the following steps: S1. Use the public semantic segmentation dataset to build similar support query sample pairs. The similar support query sample pairs include query images. , Support images and support mask ; S2, through the fusion feature extraction module The support fusion features and the query fusion features are obtained respectively, and the visual prior of the query image is obtained at the same time, wherein the fusion feature extraction module It includes a backbone network, a pair of parallel feature fusion convolutional layers, and a pair of parallel feature activation convolutional layers connected in sequence; S3, multi-scale feature extraction module Model and downsample the support fusion features and query fusion features respectively, and extract multi-scale support features and query features respectively; S4. Multi-scale prompt generation module The multi-scale support features and query features are used to generate multi-scale prompt codes, and the multi-scale prompt codes are fused from bottom to top to generate target prompt codes; wherein, the process of step S4 is as follows: S41, inputting query features and supporting features of the current scale, obtaining similarity graphs of the query features and supporting features using cosine similarity, adjusting the similarity graph using intermediate results of the next scale, aggregating the input learnable prompt vector using the adjusted similarity graph, and generating similarity prompt encoding of the current scale; S42, fusing the similar prompt code of the current scale with the prompt code and the intermediate result of the next scale to generate the prompt code of the current scale, and outputting the prompt code and the intermediate result generated by it to the previous scale to continue iteration until the target prompt code is generated; S5. Input the target hint code and the query image into the hint-type visual large model SAM to obtain the segmentation result of the query image, wherein the visual large model SAM includes an encoder and a decoder, input the query image into the encoder to obtain the query image code, and input the target hint code and the query image code into the decoder to obtain the segmentation result; construct a prediction model, which includes a fusion feature extraction module connected in sequence , Multi-scale feature extraction module , Multi-scale prompt generation module The prediction model is trained using the intermediate results and the final results to form a total loss function. During the training process, the parameters of the visual large model SAM remain frozen. S6. Input the supporting query sample pairs in the small sample data into the trained prediction model, segment the query image, and calculate the segmentation accuracy based on the prediction results.
[0039] In step S1 of this embodiment, the semantic segmentation public dataset is evenly divided into 4 groups according to categories, of which 3 groups are used as training sets and the remaining group is used as a test set. A total of 4 different combinations of training sets and test sets can be established. For the PASCAL dataset, the training set contains 15 classes, the test set contains 5 classes, and the test set contains 1000 supporting query sample pairs. Under the 1-shot small sample setting, each supporting query sample pair contains 1 query image. , 1 supporting image and 1 corresponding support mask ; Each query image and supporting images The size is ; The process of obtaining the supporting fusion features and the query fusion features and the query visual prior in step S2 of this embodiment is as follows: S21. Use the pre-trained ResNet-50 as the backbone network; select the outputs of the 2nd and 3rd layers of ResNet-50 as the middle-layer features, and the outputs of the 4th and 5th layers as the high-level features, respectively. and supporting images Enter the backbone network and obtain features:
[0040]
[0041] in, represents the pre-trained backbone network, , Respectively represent The query features and support features output by the layer, Indicates The number of channels of layer features, Indicates The height of the layer features, Indicates The width of the layer feature; Query the middle layer features respectively and support for mid-level features Input to each feature fusion convolution layer for dimension reduction and feature fusion to obtain query intermediate features and support intermediate features ; This pair of parallel feature fusion convolutional layers has the same structure, and the convolution kernel size is , the number of input channels is 1536, and the number of output channels is 256; Using high-level features and Generate visual priors by first using bilinear interpolation to support the mask The size is adjusted to support the intermediate features Consistent, using the Hadamard product method to remove high-level support features Irrelevant background features:
[0042] in, To support binary masking of images, “⊙” represents the Hadamard product; The correlation between pixels is discovered by calculating the cosine similarity of the query feature and the supporting feature at each pixel position. :
[0043] in" " represents the 3D vector inner product operation at all feature positions. represents the L2 norm; For the query feature of each position, take its maximum similarity value as the prior value of the current position, and use the minimum-maximum normalization method to normalize its value to between 0 and 1 to obtain the visual prior ; S22, using bilinear interpolation method to support mask The size is adjusted to support the intermediate features Consistently, the prototype vector of the current category is obtained using the mask average pooling operation:
[0044] in, is the prototype vector of the current category, Indicates support for intermediate features In Location The characteristic value of To support the mask In Location The annotation value of is a conditional function that takes 1 if the value in the brackets is true, and 0 otherwise; The intermediate features will be queried , prototype vector and visual priors Splicing along the channel dimension, and then passing through the query feature activation convolution layer to obtain the query fusion feature , the convolution kernel size of the query feature activation convolution layer is , the number of input channels is 514, the number of output channels is 64; intermediate features will be supported , prototype vector Splicing along the channel dimension, and then passing through the support feature activation convolution layer to obtain the support fusion feature , the convolution kernel size supporting the feature activation convolution layer is , the number of input channels is 512, and the number of output channels is 64.
[0045] In step S3 of this embodiment, the process of extracting multi-scale features by the multi-scale feature extraction module using fusion features is as follows: like Figure 2 As shown in the figure, the multi-scale feature extraction module contains three layers of attention blocks, each layer of attention blocks is connected in sequence, and the output of each layer of attention blocks is the feature information at a single scale. The outputs of the three layers of attention blocks are combined to form a multi-scale feature pyramid; each attention block includes a local attention block LocalAttnBlock and a self-attention block SelfAttnBlock connected in sequence. The formal representation of a single layer is as follows:
[0046] in, Indicates the current layer number. , It is The features of the layer output, , It is The features of the layer output, , and there is , , , Indicates The number of channels of the layer pyramid feature, , Indicates The layer pyramid features high, , Indicates The width of the layer pyramid feature, ; Each local attention block is essentially a multi-branch convolution layer, which obtains feature information under different receptive fields through convolution kernels of different sizes; the basic operation of each self-attention block is the attention mechanism :
[0047] in, is the query vector, is the key vector, is a value vector, To query features, is the key feature, is the value feature, is the corresponding projection matrix, is the value of the feature dimension; if , , If they are the same feature, it is called self-attention, if they are different, it is called mutual attention; The operation is applied to each row of the matrix to convert the attention scores into a probability distribution. , is the input matrix, For the matrix In Location The value of The query fusion features and support fusion features are input into the multi-scale feature extraction module respectively to obtain the multi-scale query feature pyramid and multi-scale support feature pyramid ,in, Represents the query feature pyramid The characteristics of the layer, Indicates support feature pyramid Characteristics of the layer.
[0048] The multi-scale prompt generation module in step S4 of this embodiment It contains 3 layers of mutual attention prompt encoding generation blocks, such as Figure 2As shown in the figure, the output of each layer is the prompt code at the current scale. Each mutual attention prompt code generation block first mines the similar information between the features of the current scale according to the cosine similarity, and then aggregates the prompt code and intermediate results of the next scale to generate the prompt code at the current scale. Specifically, the following steps are included: Step S41: Layer characteristics and , first reshape the features and shape, Reshape into , and obtain the reshaped features respectively and ; Introduce a set of learnable hint vectors , similarly, the prompt vector Reshape , respectively calculate the query vector in the attention mechanism , key vector , value vector , get the query vector , key vector Sum value vector , where the query vector From the query features , the key vector From the supporting features , value vector A learnable hint vector from a set of inputs , formally expressed as follows:
[0049] in, , , They are , , The projection matrix; calculate and The cosine similarity between them gives a similarity graph matrix :
[0050] in," " represents the vector inner product operation, represents the L2 norm of the vector, is the temperature hyperparameter, ; Along The first dimension Operation, get the normalized similarity graph matrix :
[0051] in, For the matrix In Location The value of For the matrix In Location The value of Masks will be supported using bilinear interpolation Adjust to the corresponding size and reshape to get ,use Remove the similarity values of irrelevant background positions to obtain the similarity matrix after masking the background. ; Under multi-scale conditions, the mask score output by the next layer is used as a similar prior to adjust the attention map and gradually optimize the segmentation result. The formal expression is as follows:
[0052] in, For the Layer similarity prior, obtained by resizing the logits score of the segmentation mask output by the next layer through bilinear interpolation. A logits score greater than 0 indicates foreground, and less than 0 indicates background. The larger the absolute value, the greater the probability. It is a combination of a matrix flattening operation and a dimension expansion copy operation, which first compresses the vector into a one-dimensional shape. , and then expand and replicate along the second dimension until its shape is the same as Consistency; Using the adjusted similarity matrix Aggregate the input learnable hint vector to get the first Similarity cue encoding at the layer ; Step S42: Generate a prompt code of the current scale and continue to iterate until a target prompt code is generated. The specific process is as follows: S421. If the current layer is the last layer, that is, , then directly use the similar prompt encoding at the current scale As a hint encoding at the current scale ,Right now = ; This prompt code After bilinear interpolation and resizing, it is directly input into the decoder of the visual large model SAM to obtain the intermediate result score of the last layer of segmentation , As a similarity prior input to the second layer mutual attention hint encoding generation block, the hint encoding Also enter the Iterative optimization in layers; S422. If the current layer is not the last layer, , we need to aggregate the next layer of hint codes and optimize them step by step from bottom to top. Hint encoding of layers Use bilinear interpolation to upsample to a shape consistent with the scale of the current layer. Layer Similarity Hint Encoding Spliced together, input into a fusion layer for information fusion, and obtain the prompt encoding at the current scale ; After resizing through bilinear interpolation, it is input into the decoder of the visual large model SAM to obtain the current The intermediate result score of the layer segmentation , enter it into The mutual attention hint encoding is used as a similarity prior in the generation block, hint encoding Also enter the Iterative optimization in layers; The fusion layer is essentially a multi-layer perceptron MLP, which is two linear layers connected in sequence:
[0053] in, As a nonlinear activation function, Activation function, , and is a constant; and is the weight parameter of MLP; Repeat step S422, after 3 layers of iteration, the prompt code output by the first layer is The target cue encoding is used as output.
[0054] In step S5 of this embodiment, the query image and the final hint code are input into the hint-based visual large model SAM to obtain the segmentation result, and the process of constructing the prediction model and training is as follows: S51, select a specific implementation of SAM SAMv2 as the basic model, such as Figure 1 As shown, the query image is input into the encoder of the visual large model SAM to obtain the query image code, and then the segmentation result of the query image is output according to the hint code and the query image code; the multi-scale hint code generated in step S4 is input into the decoder of the visual large model SAM together with the query image code to generate a multi-scale segmentation result ; Select As the final prediction output The remaining intermediate results are used as similarity priors to be input into the multi-scale prompt generation module middle; S52, such as Figure 3 As shown, the fusion feature extraction modules are connected in sequence , Multi-scale feature extraction module , Multi-scale prompt generation module And the visual big model SAM builds a prediction model; according to the final prediction results and intermediate results Constructing the total loss function :
[0055] in, is the total cross entropy loss, , For the mask of the query image, the cross entropy loss function , is the true label, is the predicted probability; is the distillation loss, , KL divergence loss function , is the true distribution right The probability of is the estimated distribution right The probability of For the The similarity matrix of the layer is masked average pooled and normalized along the first dimension to obtain the similarity vector; is the DICE loss of the final prediction result, , is the pixel position The probability value predicted as the result can be taken as , , is the true value of the pixel position label; , , are the weight hyperparameters of each loss, respectively. , , ; With loss function The prediction model constructed for the optimization target training is trained for 60 rounds. During the training process, the parameters of the visual large model SAM remain frozen and only the fusion feature extraction module is updated. , Multi-scale feature extraction module and multi-scale hint generation module The parameters in the training algorithm are AdamW algorithm, and the learning rate is set to , the weight decay hyperparameter is set to .
[0056] In step S6 of this embodiment, the specific process is as follows: S61. Based on the trained prediction model, the segmentation result of the query image is obtained through the forward propagation of the network; if the final segmentation mask logits score is greater than 0, it is considered to be the foreground, otherwise it is predicted to be the background; S62. Calculate the segmentation accuracy based on the prediction results, and use the average intersection over union (mIoU) as the evaluation indicator of segmentation accuracy; mIoU is the average value of the intersection over union (IoU) of all categories, and the calculation formula is as follows:
[0057]
[0058] in, For Category The number of true positive examples, i.e., correctly segmented pixels; For Category The number of false positives, i.e. pixels misclassified as class 𝑖; For Category The number of false negatives is the number of pixels that are actually of class 𝑖 but are not correctly classified; is the total number of categories.
[0059] In this example, the dataset is the public dataset PASCAL. A total of 20 classes of images are evenly divided into 4 parts, 3 of which are used as training sets and the remaining 1 is used as a test set. The small sample scene is set to 1-shot, and each support query sample pair contains 1 query image. , 1 supporting image and 1 support mask ; This example compares the proposed method with a variety of small sample semantic segmentation methods, the methods used for comparison are HSNet, BAM, HDMNet and SAM-RSP; All methods use the ResNet-50 network as the backbone network, and the segmentation results are shown in Table 1 below: Table 1. Comparison of experimental data (under 1-shot setting)
[0060] Taking mIoU as the evaluation metric, under the 1-shot setting, the average segmentation accuracy in multiple scenarios is 3.5% higher than the previous best method, and there is performance improvement under each different data partitioning.
[0061] Example 2 The overall steps of this embodiment are consistent with the overall steps in Embodiment 1.
[0062] In step S1 of this embodiment, the semantic segmentation public dataset is used to construct similar support query sample pairs. The small sample scenario adopts the 5-shot setting, that is, each support query sample pair contains 1 query image. , 5 supporting images and 5 corresponding support masks ; Other settings remain consistent with those in Example 1.
[0063] In step S2 of this embodiment, the fusion feature extraction module In the process of respectively obtaining the support fusion features and the query fusion features and simultaneously obtaining the visual prior of the query image, compared with Example 1, the visual prior obtained in this embodiment is and support fusion features The number is 5 times that in Example 1; for visual prior , directly take the average of the five visual priors as the final visual prior, which supports the fusion feature No other processing is performed for the time being, and the other processes remain consistent with the process in step S2 in Example 1.
[0064] The process of extracting multi-scale features by the multi-scale feature extraction module using fusion features in step S3 of this embodiment refers to step S3 in embodiment 1, and the number of multi-scale supporting features obtained is 5 times that in embodiment 1.
[0065] The multi-scale prompt generation module in step S4 of this embodiment It contains 3 layers of mutual attention hint coding generation blocks. The output of each layer is the hint coding at the current scale. Each mutual attention hint coding generation block first mines the similar information between the features of the current scale according to the cosine similarity, and then aggregates the hint coding and intermediate results of the next scale to generate the hint coding at the current scale. Specifically, it includes the following steps: S41. The specific steps are consistent with step S41 in Example 1. The difference is that since the number of multi-scale supporting features obtained in this embodiment is 5 times that of Example 1, the similarity prompt code generated at a single scale is The number of is also 5 times that of Example 1, and the 5 similar prompt codes are calculated by averaging. The average value of is used as the final similarity hint encoding at the current scale. ; S42, refer to step S42 in embodiment 1; In step S5 of this embodiment, the query image and the final hint code are input into the hint-based visual large model SAM to obtain the segmentation result. For specific steps, refer to step S5 in embodiment 1.
[0066] In step S6 of this embodiment, the supporting query sample pairs in the small sample data are input into the trained prediction model, the query image is segmented, and the segmentation accuracy is calculated according to the prediction results. The specific steps refer to step S6 in embodiment 1.
[0067] In this example, the dataset is the public dataset PASCAL. A total of 20 classes of images are evenly divided into 4 parts, 3 of which are used as training sets and the remaining 1 is used as a test set. The small sample scenario is set to 5-shot, and each support query sample pair contains 1 query image. , 5 supporting images and 5 corresponding support masks ; This embodiment compares the proposed method with a variety of small sample semantic segmentation methods, and the methods used for comparison are HSNet, BAM, HDMNet and SAM-RSP; All methods use the ResNet-50 network as the backbone network. The segmentation results of the present invention and other methods in the 5-shot setting are compared as shown in Table 2 below: Table 2. Comparison of experimental data (under 5-shot setting)
[0068] Taking mIoU as the evaluation metric, under the 5-shot setting, the average segmentation accuracy in multiple scenarios is 2.0% higher than the previous best method, and there is performance improvement under each different data partitioning.
[0069] The results of Example 1 and Example 2 show that the present invention can be used in a small sample semantic segmentation visual task system. Under the condition of having only a small number of annotated support samples, it can segment out new classes of objects that have not been seen in the model training process, which can greatly reduce the labeling cost of new classes of target objects. The present invention can further mine rich semantic information from multi-scale feature information in a variety of scenarios and different task settings. The generated target hint coding can more accurately point to the analog object to be segmented in the current scenario, and the final segmentation result can also achieve a higher average accuracy rate. Therefore, the target hint coding generated by continuously iteratively optimizing the use of multi-scale information proposed by the present invention is more robust.
[0070] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited to the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously.
[0071] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0072] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A small sample semantic segmentation method based on multi-scale cues from a large visual model, characterized in that: The small sample semantic segmentation method comprises the following steps: S1. Use the public semantic segmentation dataset to build similar support query sample pairs. The similar support query sample pairs include query images. , Support images and support mask ; S2, through the fusion feature extraction module The support fusion features and the query fusion features are obtained respectively, and the visual prior of the query image is obtained at the same time, wherein the fusion feature extraction module It includes a backbone network, a pair of parallel feature fusion convolutional layers, and a pair of parallel feature activation convolutional layers connected in sequence; S3, multi-scale feature extraction module Model and downsample the support fusion features and query fusion features respectively, and extract multi-scale support features and query features respectively; S4, Multi-scale Hint Generation Module The multi-scale support features and query features are used to generate multi-scale prompt codes, and the multi-scale prompt codes are fused from bottom to top to generate target prompt codes; wherein, the process of step S4 is as follows: S41, inputting query features and supporting features of the current scale, obtaining similarity graphs of the query features and supporting features using cosine similarity, adjusting the similarity graph using intermediate results of the next scale, aggregating the input learnable prompt vector using the adjusted similarity graph, and generating similarity prompt encoding of the current scale; S42, fusing the similar prompt code of the current scale with the prompt code and the intermediate result of the next scale to generate the prompt code of the current scale, and outputting the prompt code and the intermediate result generated by it to the previous scale to continue iteration until the target prompt code is generated; S5. Input the target hint code and the query image into the hint-type visual large model SAM to obtain the segmentation result of the query image, wherein the visual large model SAM includes an encoder and a decoder, input the query image into the encoder to obtain the query image code, and input the target hint code and the query image code into the decoder to obtain the segmentation result; construct a prediction model, which includes a fusion feature extraction module connected in sequence , Multi-scale feature extraction module , Multi-scale prompt generation module The prediction model is trained using the intermediate results and the final results to form a total loss function. During the training process, the parameters of the visual large model SAM remain frozen. S6. Input the supporting query sample pairs in the small sample data into the trained prediction model, segment the query image, and calculate the segmentation accuracy based on the prediction results.
2. According to claim 1, a small sample semantic segmentation method based on multi-scale cues of a large visual model is characterized in that: The process of constructing the same type of supporting query sample pairs in step S1 is as follows: The public datasets for semantic segmentation are divided into four groups by category, of which three training sets are used for training and the remaining one test set is used for testing. In each set of data, the training set and the test set contain a support set and a query set respectively. The support set consists of support images and their corresponding densely annotated masks, while the query set consists of single images of the same type and their masks. The densely annotated masks of the support images in the support set are used as prior knowledge. During the training process, the prediction network is supervised by the densely annotated masks of the query set images to optimize the network parameters. During the testing phase, densely annotated masks of query images are used as criteria to evaluate model performance.
3. According to claim 1, a small sample semantic segmentation method based on multi-scale cues of a large visual model is characterized in that: The process of obtaining the supporting fusion features and querying the fusion features and querying the visual prior in step S2 is as follows: S21. Use the pre-trained ResNet-50 as the backbone network, select the outputs of the 2nd and 3rd layers in ResNet-50 as the middle-layer features, and the outputs of the 4th and 5th layers as the high-level features, and respectively convert the query image and supporting images Input, get features: in, represents the pre-trained backbone network, , Respectively represent The query features and support features output by the layer, Indicates The number of channels of layer features, Indicates The height of the layer features, Indicates The width of the layer feature; Query the middle layer features respectively and support for mid-level features Input to each feature fusion convolution layer for dimension reduction and feature fusion to obtain query intermediate features and support intermediate features ; Using high-level features and Generating visual priors using support masks Remove high-level support features The irrelevant background features in the image are used to discover the correlation between pixels by calculating the cosine similarity of each pixel position between the query feature and the supporting feature. : in" " represents the 3D vector inner product operation at all feature positions. represents the L2 norm; For the query feature of each position, take its maximum similarity value as the prior value of the current position, and use the minimum-maximum normalization method to normalize its value to between 0 and 1 to obtain the visual prior ; S22, using bilinear interpolation method to support mask The size is adjusted to support the intermediate features Consistently, the prototype vector of the current category is obtained using the mask average pooling operation: in, is the prototype vector of the current category, Indicates support for intermediate features In Location The characteristic value of To support the mask In Location The annotation value of is a conditional function that takes 1 if the value in the brackets is true, and 0 otherwise; The intermediate features will be queried , prototype vector and visual priors Splicing along the channel dimension, and then passing through the query feature activation convolution layer to obtain the query fusion feature ; Will support intermediate features , prototype vector Splicing along the channel dimension, and then passing through the support feature activation convolution layer to obtain the support fusion feature .
4. According to claim 1, a small sample semantic segmentation method based on multi-scale cues of a large visual model is characterized in that: The process of extracting multi-scale features by the multi-scale feature extraction module using fusion features in step S3 is as follows: The multi-scale feature extraction module includes Each layer of attention blocks is connected in sequence, and the output of each layer of attention blocks is the feature information at a single scale. The outputs of the layer attention blocks are combined to form a multi-scale feature pyramid; each attention block includes a local attention block LocalAttnBlock and a self-attention block SelfAttnBlock connected in sequence. The formal representation of a single layer is as follows: in, Indicates the current layer number. It is The features of the layer output, , It is The features of the layer output, , and there is , , , Indicates The number of channels of the layer pyramid feature, Indicates The layer pyramid features high, Indicates The width of the layer pyramid feature; Each local attention block is essentially a multi-branch convolution layer, which obtains feature information under different receptive fields through convolution kernels of different sizes; the basic operation of each self-attention block is the attention mechanism : in, is the query vector, is the key vector, is a value vector, To query features, is the key feature, is the value feature, is the corresponding projection matrix, is the value of the feature dimension; if , , If they are the same feature, it is called self-attention, if they are different, it is called mutual attention; The operation is applied to each row of the matrix to convert the attention scores into a probability distribution. , is the input matrix, For the matrix In Location The value of The query fusion features and support fusion features are input into the multi-scale feature extraction module respectively to obtain the multi-scale query feature pyramid and multi-scale support feature pyramid ,in, Represents the query feature pyramid The characteristics of the layer, Indicates support feature pyramid Characteristics of the layer.
5. According to claim 4, a small sample semantic segmentation method based on multi-scale cues of a large visual model is characterized in that: The multi-scale prompt generation module in step S4 Include The output of each layer is the hint code at the current scale. Each mutual attention hint code generation block first mines the similarity information between the features of the current scale according to the cosine similarity, and then aggregates the hint code and intermediate results of the next scale to generate the hint code at the current scale.
6. The small sample semantic segmentation method based on multi-scale cues of a large visual model according to claim 5, characterized in that: The process of generating the similarity hint code of the current scale in step S41 is as follows: For Layer characteristics and , first reshape the features and shape, Reshape into , and obtain the reshaped features respectively and , input the learnable hint vector , similarly, the prompt vector Reshape , respectively calculate the query vector in the attention mechanism , key vector , value vector , get the query vector , key vector Sum value vector , where the query vector From the query features , the key vector From the supporting features , value vector A learnable hint vector from a set of inputs , formally expressed as follows: in, , , They are , , The projection matrix; In order to obtain the similarity relationship between each pixel position, calculate and The cosine similarity between them gives a similarity graph matrix : in," " represents the vector inner product operation, represents the L2 norm of the vector, is the temperature hyperparameter; Along The first dimension Operation, get the normalized similarity graph matrix : in, For the matrix In Location The value of For the matrix In Location The value of Masks will be supported using bilinear interpolation Adjust to the corresponding size and reshape to get ,use Remove the similarity values of irrelevant background positions to obtain the similarity matrix after masking the background. ; Under multi-scale conditions, the mask score output by the next layer is used as a similar prior to adjust the attention map and gradually optimize the segmentation result. The formal expression is as follows: in, For the Layer similarity prior, obtained by resizing the logits score of the segmentation mask output by the next layer through bilinear interpolation. A logits score greater than 0 indicates foreground, and less than 0 indicates background. The larger the absolute value, the greater the probability. It is a combination of a matrix flattening operation and a dimension expansion copy operation, which first compresses the vector into a one-dimensional shape. , and then expand and copy along the second dimension until the shape is the same as Consistency; Using the adjusted similarity matrix Aggregate the input learnable hint vector to get the first Similarity cue encoding at the layer .
7. The small sample semantic segmentation method based on multi-scale cues of a large visual model according to claim 6, characterized in that: The process of generating the prompt code of the current scale and continuing to iterate to generate the target prompt code in step S42 is as follows: S421. If the current layer is the last layer, that is, , then directly use the similar prompt encoding at the current scale As a hint encoding at the current scale ,Right now = ; This prompt code After bilinear interpolation and resizing, it is directly input into the decoder of the visual large model SAM to obtain the intermediate result score of the last layer of segmentation , As a similarity prior input to the In the layer mutual attention hint encoding generation block, the hint encoding Also enter the Iterative optimization in layers; S422. If the current layer is not the last layer, , we need to aggregate the next layer of hint codes and optimize them step by step from bottom to top. Hint encoding of layers Use bilinear interpolation to upsample to a shape consistent with the scale of the current layer. Layer Similarity Hint Encoding Spliced together, input into a fusion layer for information fusion, and obtain the prompt encoding at the current scale ; After resizing through bilinear interpolation, it is input into the decoder of the visual large model SAM to obtain the current The intermediate result score of the layer segmentation , enter it into The mutual attention hint encoding is used as a similarity prior in the generation block, hint encoding Also enter the Iterative optimization in layers; Repeat step S422, after After the iteration of the layer, the prompt encoding of the first layer output The target cue encoding is used as output.
8. The small sample semantic segmentation method based on multi-scale cues of a large visual model according to claim 7, characterized in that: In step S5, the query image and the final hint code are input into the hint-based visual large model SAM to obtain the segmentation result, and the process of constructing the prediction model and training is as follows: S51, inputting the query image into the encoder of the visual large model SAM to obtain the query image code, and then outputting the segmentation result of the query image according to the prompt code and the query image code; The multi-scale hint code generated in step S4 is input into the decoder of the visual large model SAM together with the query image code to generate a multi-scale segmentation result ; Select As the final prediction output , the rest of the intermediate results Used as similarity priors to feed into the multi-scale cue generation module middle; S52, successive fusion feature extraction modules , Multi-scale feature extraction module , Multi-scale prompt generation module And the visual big model SAM builds a prediction model; according to the final prediction results and intermediate results Constructing the total loss function , with the loss function The prediction model constructed for optimization goal training, loss function The calculation method is as follows: in, is the total cross entropy loss, , For the mask of the query image, the cross entropy loss function , is the true label, is the predicted probability; is the distillation loss of the similarity matrix between layers of the multi-scale hint generation module, , KL divergence loss function , is the true distribution right The probability of is the estimated distribution right The probability of For the The similarity matrix of the layer is masked average pooled and normalized along the first dimension to obtain the similarity vector; is the DICE loss of the final prediction result, , is the pixel position The probability value predicted as the outcome, is the true value of the pixel position label; , , are the weight hyperparameters of each loss respectively.
9. The small sample semantic segmentation method based on multi-scale cues of a large visual model according to claim 1, characterized in that: The process in step S6 is as follows: S61. Based on the trained prediction model, the segmentation result of the query image is obtained through the forward propagation of the network; if the final segmentation mask logits score is greater than 0, it is considered to be the foreground, otherwise it is predicted to be the background; S62. Calculate the segmentation accuracy based on the prediction results, and use the average intersection over union (mIoU) as an evaluation indicator of the segmentation accuracy.
Citation Information
Patent Citations
Small sample image segmentation method and device, terminal equipment and storage medium
CN116030250A
SAM-based small-sample industrial anomaly segmentation deep learning method and device, and terminal
CN117877030A
Small sample image classification method and system based on multi-scale cross-modal prompt enhancement
CN119313966A
Cited By
Underwater hull fouling marking optimization method based on multi-modal large model guidance
CN120689737A
Unified medical image segmentation method based on context hierarchical guidance
CN121073995A
Automatic eye muscle segmentation method and system based on large model fine tuning
CN121640057A
Alloy microstructure image segmentation method and device, computer equipment and medium
CN122618240A