A semantic segmentation method for small sample images based on feature separation and recombination

By using feature separation and recombination methods in semantic segmentation of small sample images, using background information and target information that support the picture, high robustness and high reliability semantic feature representations are generated, which solves the problem that existing methods are difficult to accurately segment unknown new class targets, and significantly improves segmentation performance.

CN116805368BActive Publication Date: 2025-06-06LANZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310858480.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-13
Publication Date
2025-06-06
Estimated Expiration
2043-07-13

AI Technical Summary

Technical Problem

The existing small sample image semantic segmentation method only uses the deep semantic features of the target prospect, making it difficult to fully mine and query the information in the picture, resulting in the robustness and reliability of the specific semantic representation of the learned class, and it is difficult to accurately segment out unknown new class targets.

Method used

Through the method of feature separation and recombination, the background information and target information supporting pictures are used to separate the support features into target foreground features and background features, the query features are separated into pseudo-target foreground features and pseudo-background features, and feature recombination is performed through self-attention and cross-attention to generate a semantic feature representation with high robustness and high reliability.

Benefits of technology

It improves the robustness and reliability of class-specific semantic representation, improves the performance of semantic segmentation of small sample images, and can more accurately segment out unknown new class targets in query images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116805368B_ABST
    Figure CN116805368B_ABST
Patent Text Reader

Abstract

The present invention discloses a small sample image semantic segmentation method based on feature separation and recombination, comprising the steps of: feature extraction, mapping support images and query images to deep feature space using a backbone network; feature separation, separating query features into pseudo target foreground features and pseudo background features of equal size; feature recombination, semantically reorganizing the background features of the support branch and the query branch; prototype learning, generating feature prototypes corresponding to the foreground and background regions using global average pooling; parameter-free measurement, guiding the segmentation of targets in the query image and the support image by calculating the similarity between the query feature, the support feature and the feature prototype. The present invention can not only realize the feature separation of the query feature, but also realize the recombination of the separation feature and the support feature, thereby improving the reliability of feature expression, realizing the semantic segmentation of the support and query dual-branch input images, and providing a new idea for segmentation in the fields of natural scenes, medical images, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a small sample image semantic segmentation method based on feature separation and recombination. Background Art

[0002] In recent years, deep learning has made breakthrough progress in the fields of vision, text, and speech, largely due to a large number of labeled training samples. However, data labeling is time-consuming and laborious, especially for semantic segmentation tasks, which require a large number of dense pixel-level annotations and are extremely costly. Although the use of weak supervision methods can reduce the high cost of collecting labeled data to a certain extent, the generalization performance of such methods for unknown classes is not high. Small sample learning can learn matching rules from a small number of samples with labeled information and quickly generalize to the segmentation task of unknown new classes, which can alleviate the problem of high data labeling costs to a certain extent.

[0003] Most existing small sample image semantic segmentation methods adopt a dual-branch network structure of support branch and query branch, which uses the support branch and query branch to map the support image and query image to the deep feature space respectively, and establishes rules to guide the segmentation of the query image by deeply learning the deep semantic features of each image in the support set. Although the segmentation performance of existing deep learning models is constantly improving, this type of method only uses the deep semantic features of the target foreground to guide the segmentation of the query image, and does not fully explore the information contained in the query image, making the learned class-specific semantic representations not robust and reliable, and it is difficult to accurately segment unknown new class targets in the query image. Summary of the invention

[0004] In order to overcome the shortcomings of the existing basis, the present invention proposes a small sample image semantic segmentation method which fully utilizes the background information of the supporting image and further mines the target information to learn highly robust and reliable semantic feature representation, and further improve the performance of small sample image semantic segmentation.

[0005] The technical solution of the present invention is implemented as follows: a small sample image semantic segmentation method based on feature separation and recombination comprises the following steps:

[0006] S1, feature extraction, using the backbone network to map support images and query images into deep feature space;

[0007] S2, feature separation, uses the true mask of the support image to separate the support features into support target foreground features and support background features, and separates the query features into pseudo target foreground features and pseudo background features of equal size;

[0008] S3, feature reorganization, by establishing the self-attention of the support target foreground feature, the self-attention of the pseudo target foreground feature, the cross-attention of the support target foreground feature and the self-attention of the pseudo target foreground feature, the target foreground features of the support branch and the query branch are semantically reorganized, and by establishing the cross-attention of the support background feature and the query pseudo background feature, the background features of the support branch and the query branch are semantically reorganized;

[0009] S4, prototype learning, uses global average pooling on the reorganized foreground feature and background feature maps to generate feature prototypes corresponding to the foreground and background regions;

[0010] S5, a parameter-free metric, guides the segmentation of objects in query and support images by calculating the similarity between query features, support features, and feature prototypes.

[0011] Furthermore, the backbone network described in step S1 uses pre-trained VGG-16, ResNet-50 and ResNet-101 networks with shared weights to map support images and query images to a deep feature space.

[0012] Furthermore, the query feature separation described in step S2 is to use the original query feature containing the foreground and the background as the initial feature of the pseudo target foreground feature and the pseudo background feature, and establish interactions with the supporting target foreground feature and the supporting background feature respectively, so as to separate the original query feature containing the foreground and the background into the pseudo target foreground feature and the pseudo background feature.

[0013] Furthermore, the feature recombination described in step S3 comprises the steps of:

[0014] S31, establish the self-attention of the target foreground feature, the self-attention calculation of the target foreground feature is shown in formula (1),

[0015]

[0016] S32, establish the self-attention of the pseudo target foreground feature, the pseudo target foreground feature self-attention calculation is shown in formula (2),

[0017]

[0018] S33, establish cross attention that supports the target foreground feature self-attention and the pseudo target foreground feature self-attention, and semantically reorganize the target foreground features of the support branch and the query branch. The foreground cross attention calculation is shown in formula (3):

[0019]

[0020] S34, establish cross attention of support background features and query pseudo background features, and semantically reorganize the background features of the support branch and the query branch. The background cross attention calculation is shown in formula (4):

[0021]

[0022] Among them, F sup_fg Represents the self-attention of the target foreground features; F que_fg represents the pseudo target foreground feature self-attention; F cs_fg Ff represents the cross attention of target foreground feature self-attention and pseudo target foreground feature self-attention; s Indicates the foreground feature that supports the target; Ff q Indicates the pseudo target foreground feature of the query; Fb s Indicates support for background features; Fb q represents the query pseudo-background feature; d represents the feature dimension.

[0023] Furthermore, the prototype learning described in step S4 includes the steps of:

[0024] S41, using global average pooling on the reorganized foreground feature map to generate a feature prototype corresponding to the foreground area, the foreground feature prototype is calculated as shown in formula (5),

[0025]

[0026] S42: Generate a feature prototype corresponding to the background area using global average pooling on the reorganized background feature map. The background feature prototype is calculated as shown in formula (6):

[0027]

[0028] S43: Concatenate the foreground feature prototype and the background feature prototype to obtain a prototype set P = {p_fg}∪{p_bg}.

[0029] Among them, p_fg represents the foreground feature prototype; p_bg represents the background feature prototype; K represents the total number of supported images; represents the truth function; M k represents the mask corresponding to the k-th support image; c represents the class label corresponding to the k-th support image.

[0030] Furthermore, the parameter-free measurement described in step S5 includes the steps of:

[0031] S51: Calculate the feature Fq of the query image mapped to the deep feature space (x, y) (x,y) Cosine similarity S with each prototype q (x,y), the similarity calculation is shown in formula (7),

[0032]

[0033] S52: Calculate the feature Fs that supports the image mapping to the deep feature space (x, y) (x,y) Cosine similarity S with each prototype s (x,y) , the similarity calculation is shown in formula (8),

[0034]

[0035] S53: Use the argmax(·) function to select the maximum similarity value between the feature at each position (x, y) and the prototype set P, and concatenate the maximum similarity values ​​at all positions, and use bilinear interpolation to generate masks of the query image and the support image. The generated mask can be expressed as formula (9):

[0036] M′=∑ x,y BIL(argmax(S (x,y) )) (9)

[0037] Among them, Fq (x,y) Fs represents the feature representation of the query image mapped to the deep feature space (x, y); (x,y) Indicates the feature representation that supports the image mapping to the deep feature space (x, y); M′ represents the generated mask; BIL represents bilinear interpolation;

[0038] S54: Calculate the cross entropy loss ls between the predicted support image mask and the true support image mask s→s , calculate the cross entropy loss ls between the predicted query image mask and the true query image mask s→q , the network model is optimized end-to-end using the dual loss of the support branch and the query branch. The dual loss calculation is shown in formula (10):

[0039]

[0040] Among them, h and w are the length and width of the supported image; α is a learnable parameter; is the real support mask, is the support mask for the prediction; is the actual query mask, which is only used in the training phase. The query mask for the prediction.

[0041] Further or more further, the supporting background feature described in step S3 is the average feature after the supporting background feature map is divided into n = (h / σ) × (w / σ) patch blocks using VisionTransformer, where n represents the number of divided patch blocks; h and w are the length and width of the supporting image, and σ represents the size of the convolution kernel.

[0042] The beneficial effect of the present invention is that, compared with the prior art, the present invention can not only realize the feature separation of query features, but also realize the recombination of separation features and support features, thereby improving the robustness and reliability of class-specific semantic representation. In addition, the semantic segmentation method disclosed in the present invention can also realize the semantic segmentation of support and query dual-branch input images, providing a new idea for segmentation in the fields of natural scenes, medical images, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 This is a flow chart of a small sample image semantic segmentation method based on feature separation and recombination of the present invention;

[0044] Figure 2 This is a prototype learning flowchart of a small sample image semantic segmentation method based on feature separation and recombination of the present invention. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] The process of the semantic segmentation method of a small sample image based on feature separation and recombination of the present invention is as follows: Figure 1As shown, the technical solution of the present invention is: using the backbone network to map the support image and the query image to the deep feature space, and using the real mask of the support image to separate the support features into the support target foreground features and the support background features, and to separate the query features into pseudo target foreground features and pseudo background features of equal size. Secondly, by establishing the self-attention of the support target foreground features, the self-attention of the pseudo target foreground features, the cross-attention of the support target foreground features and the self-attention of the pseudo target foreground features, the semantic reorganization of the target foreground features of the support branch and the query branch is carried out, and by establishing the cross-attention of the support background features and the query pseudo background features, the semantic reorganization of the background features of the support branch and the query branch is carried out. Thirdly, the feature prototypes corresponding to the foreground and background regions are generated by global average pooling on the reorganized foreground feature and background feature map. Finally, by calculating the similarity between the query feature, the support feature and the feature prototype, the target in the query image and the support image is guided to be segmented.

[0047] The present invention provides a small sample image semantic segmentation method based on feature separation and recombination, which comprises the following steps:

[0048] S1: Feature extraction, using the backbone network to map the support image and the query image to the deep feature space, wherein the backbone network uses a pre-trained VGG-16, ResNet-50 and ResNet-101 network with shared weights to map the support image and the query image to the deep feature space.

[0049] S2: Feature separation, using the real mask of the support image to separate the support features into support target foreground features and support background features, and separate the query features into pseudo target foreground features and pseudo background features of equal size, wherein the query feature separation is to use the original query features containing the foreground and background as the initial features of the pseudo target foreground features and the pseudo background features, and establish interactions with the support target foreground features and the support background features respectively, and separate the original query features containing the foreground and the background into pseudo target foreground features and pseudo background features.

[0050] S3: Feature reorganization, by establishing the self-attention of the support target foreground features, the self-attention of the pseudo target foreground features, and the cross-attention of the support target foreground features and the pseudo target foreground features, the target foreground features of the support branch and the query branch are semantically reorganized; by establishing the cross-attention of the support background features and the query pseudo background features, the background features of the support branch and the query branch are semantically reorganized.

[0051] The main process for the above feature recombination step is:

[0052] S31: Establishing the self-attention of the target foreground feature. The self-attention calculation of the target foreground feature is shown in formula (1).

[0053]

[0054] S32: Establishing the self-attention of the pseudo target foreground feature. The pseudo target foreground feature self-attention calculation is shown in formula (2):

[0055]

[0056] S33: Establish cross attention that supports the target foreground feature self-attention and the pseudo target foreground feature self-attention, and semantically reorganize the target foreground features of the support branch and the query branch. The foreground cross attention calculation is shown in formula (3):

[0057]

[0058] S34: Establish cross attention of support background features and query pseudo background features, and semantically reorganize the background features of the support branch and the query branch, wherein the support background feature is the average feature after the support background feature map is divided into n = (h / σ) × (w / σ) patches using Vision Transformer, where n represents the number of divided patches; h and w are the length and width of the support image, and σ represents the size of the convolution kernel. The background cross attention calculation is shown in formula (4),

[0059]

[0060] Among them, F sup_fg Represents the self-attention of the target foreground features; F que_fg represents the pseudo target foreground feature self-attention; F cs_fg Ffs represents the cross attention of target foreground feature self-attention and pseudo target foreground feature self-attention; s Indicates the foreground feature that supports the target; Ff q Indicates the pseudo target foreground feature of the query; Fb s Indicates support for background features; Fb q represents the query pseudo-background feature; d represents the feature dimension.

[0061] S4: Prototype learning, using global average pooling on the reorganized foreground feature and background feature maps to generate feature prototypes corresponding to the foreground and background regions. The process of prototype learning is as follows: Figure 2 shown.

[0062] The main process for the aforementioned prototype learning steps is:

[0063] S41: Generate a feature prototype corresponding to the foreground region using global average pooling on the reorganized foreground feature map. The foreground feature prototype is calculated as shown in formula (5):

[0064]

[0065] S42: Generate a feature prototype corresponding to the background area using global average pooling on the reorganized background feature map. The background feature prototype is calculated as shown in formula (6):

[0066]

[0067] S43: Concatenate the foreground feature prototype and the background feature prototype to obtain a prototype set P = {p_fg}∪{p_bg}.

[0068] Among them, p_fg represents the foreground feature prototype; p_bg represents the background feature prototype; K represents the total number of supported images; represents the truth function; M k represents the mask corresponding to the k-th support image; c represents the class label corresponding to the k-th support image.

[0069] S5: A parameter-free metric that guides the segmentation of objects in query and support images by calculating the similarity between query features, support features, and feature prototypes.

[0070] The main process for the aforementioned parameter-free measurement step is:

[0071] S51: Calculate the feature Fq of the query image mapped to the deep feature space (x, y) (x,y) Cosine similarity S with each prototype q (x,y) , the similarity calculation is shown in formula (7),

[0072]

[0073] S52: Calculate the feature Fs that supports the image mapping to the deep feature space (x, y) (x,y) Cosine similarity S with each prototype s (x,y) , the similarity calculation is shown in formula (8),

[0074]

[0075] S53: Use the argmax(·) function to select the maximum similarity value between the feature at each position (x, y) and the prototype set P, and concatenate the maximum similarity values ​​at all positions, and use bilinear interpolation to generate masks of the query image and the support image. The generated mask can be expressed as formula (9):

[0076] M′=∑ x,y BIL(argmax(S (x,y))) (9)

[0077] Among them, Fq (x,y) Fs represents the feature representation of the query image mapped to the deep feature space (x, y); (x,y) It represents the feature representation that supports mapping the image to the deep feature space (x, y); M′ represents the generated mask; BIL represents bilinear interpolation.

[0078] S54: Calculate the cross entropy loss ls between the predicted support image mask and the true support image mask s→s , calculate the cross entropy loss ls between the predicted query image mask and the true query image mask s→q , the network model is optimized end-to-end using the dual loss of the support branch and the query branch. The dual loss calculation is shown in formula (10):

[0079]

[0080] Among them, h and w are the length and width of the supported image; α is a learnable parameter; is the real support mask, is the support mask for the prediction; is the actual query mask, which is only used in the training phase. The query mask for the prediction.

[0081] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A semantic segmentation method for small sample images based on feature separation and recombination, It is characterized in that Includes steps: S1, feature extraction, using the backbone network to map support images and query images into deep feature space; S2, feature separation, uses the true mask of the support image to separate the support features into support target foreground features and support background features, and separates the query features into pseudo target foreground features and pseudo background features of equal size; S3, feature reorganization, by establishing the self-attention of the support target foreground feature, the self-attention of the pseudo target foreground feature, the cross-attention of the support target foreground feature and the self-attention of the pseudo target foreground feature, the target foreground features of the support branch and the query branch are semantically reorganized, and by establishing the cross-attention of the support background feature and the query pseudo background feature, the background features of the support branch and the query branch are semantically reorganized; S4, prototype learning, using global average pooling on the reorganized foreground feature and background feature maps to generate feature prototypes corresponding to the foreground and background regions, the prototype learning includes the following steps: S41, using global average pooling on the reorganized foreground feature map to generate a feature prototype corresponding to the foreground area, the foreground feature prototype is calculated as shown in formula (5), S42: Generate a feature prototype corresponding to the background area using global average pooling on the reorganized background feature map. The background feature prototype is calculated as shown in formula (6): S43: concatenate the foreground feature prototype and the background feature prototype to obtain a prototype set P = {p_fg}∪{p_bg}; Among them, p_fg represents the foreground feature prototype; p_bg represents the background feature prototype; K represents the total number of supported images; represents the truth function; M k represents the mask corresponding to the k-th support image; c represents the class label corresponding to the k-th support image S5, parameter-free measurement, by calculating the similarity between the query feature, the support feature and the feature prototype, guiding the segmentation of the objects in the query image and the support image, the parameter-free measurement includes the steps of: S51: Calculate the feature Fq of the query image mapped to the deep feature space (x, y) (x,y) Cosine similarity with each prototype S52: Calculate the feature Fs that supports the image mapping to the deep feature space (x, y) (x,y) Cosine similarity with each prototype S53: Use the argmax(·) function to select the maximum similarity value between the feature at each position (x, y) and the prototype set P, and concatenate the maximum similarity values ​​at all positions, and use bilinear interpolation to generate masks of the query image and the support image. The generated mask can be expressed as formula (9): M′=∑ x,y BIL(argmax(S (x,y) )) (9) Where M′ represents the generated mask; BIL represents bilinear interpolation; S54: Calculate the cross entropy loss ls between the predicted support image mask and the true support image mask s→s , calculated as shown in formula (10), Among them, h and w are the length and width of the supported image; α is a learnable parameter; is the real support mask, is the support mask for the prediction; is the actual query mask, which is only used in the training phase. The query mask for the prediction.

2. The small sample image semantic segmentation method based on feature separation and recombination as claimed in claim 1, It is characterized in that The backbone network described in step S1 uses pre-trained VGG-16, ResNet-50 and ResNet-101 networks with shared weights to map support images and query images to a deep feature space.

3. The small sample image semantic segmentation method based on feature separation and recombination as claimed in claim 1, It is characterized in that The query feature separation described in step S2 is to use the original query features containing the foreground and the background as the initial features of the pseudo target foreground features and the pseudo background features, and establish interactions with the supporting target foreground features and the supporting background features respectively, so as to separate the original query features containing the foreground and the background into the pseudo target foreground features and the pseudo background features.

4. The small sample image semantic segmentation method based on feature separation and recombination as claimed in claim 1, It is characterized in that The feature recombination described in step S3 comprises the steps of: S31, establish the self-attention of the target foreground feature, the self-attention calculation of the target foreground feature is shown in formula (1), S32, establish the self-attention of the pseudo target foreground feature, the pseudo target foreground feature self-attention calculation is shown in formula (2), S33, establish cross attention that supports the target foreground feature self-attention and the pseudo target foreground feature self-attention, and semantically reorganize the target foreground features of the support branch and the query branch. The foreground cross attention calculation is shown in formula (3): S34, establish cross attention of support background features and query pseudo background features, and semantically reorganize the background features of the support branch and the query branch. The background cross attention calculation is shown in formula (4): in, / sup_fg Represents the self-attention of the target foreground features; F que_fg represents the pseudo target foreground feature self-attention; F cs_fg Ff represents the cross attention of target foreground feature self-attention and pseudo target foreground feature self-attention; s Indicates the foreground feature that supports the target; Ff q Fb represents the pseudo target foreground feature of the query; s Indicates support for background features; Fb q represents the query pseudo-background feature; d represents the feature dimension.

5. The small sample image semantic segmentation method based on feature separation and recombination as claimed in claim 1, It is characterized in that The support background feature described in step S3 is the average feature after the support background feature map is divided into n = (h / σ) × (w / σ) patch blocks using Vision Transformer, where n represents the number of divided patch blocks; h and w are the length and width of the support image, and σ represents the size of the convolution kernel.