Few-shot Segmentation Method Based on Generalized Features and Region Proposal Loss Function
In the field of small sample image segmentation, a small sample segmentation method based on generalized features and region suggestion loss function is adopted, and the problem of unsatisfactory segmentation effect when the number of samples is small is solved, and efficient image segmentation under low data volume is achieved.
Patent Information
- Application Number
- CN202211390317.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-03
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-11-03
AI Technical Summary
In the field of small sample image segmentation, the prior art is difficult to achieve ideal segmentation effects in the field with small samples, and overfitting problems are prone to occur.
A small sample segmentation method based on generalization features and regional suggestion loss function is proposed. Through the training of a variety of data augmentation and deformable convolutional network, the generalization ability of the model is improved. The channel attention and spatial attention modules are used to extract depth features, and the regional suggestion loss function is introduced for positional positioning, and finally the model fine-tuning is performed under low data volume.
With fewer samples training, rapid convergence and better segmentation performance are achieved, the generalization ability of the model is improved, the data overfitting problem is solved, and the generalization of the model is improved.
Smart Images

Figure CN115661163B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of few-shot image segmentation and computer vision, and particularly relates to a few-shot segmentation method based on generalization features and a region proposal loss function. Background Art
[0002] In recent years, deep convolutional neural networks have achieved great success in the field of computer vision, and image segmentation methods based on deep learning can also have good effects on public datasets through more accurate feature extraction. Moreover, as the accuracy of image segmentation continues to improve, image segmentation in various refined branches has gradually matured, and there are various deep convolutional networks for image segmentation in different fields. In such a good situation, the rapid development of image segmentation depends to a large extent on the training of a large number of datasets and labels. However, for some fields with a small number of samples and corresponding labels, the model is often prone to overfitting, resulting in unsatisfactory segmentation effects. Therefore, the field of few-shot image segmentation has emerged.
[0003] Few-shot image segmentation relies on learning the generalization features of images based on a large number of base class datasets, so that fine-tuning in fields with fewer samples can achieve good effects. However, the current metrics in this direction still need to be improved. Summary of the Invention
[0004] Image segmentation, as one of the more commonly used algorithms in computer vision, requires a large number of annotated images and a large amount of time, labor, and cost. In the case of low data volume, due to the diversity of data, it is difficult for only a small amount of data to show good performance. To solve the above problems, the present invention proposes a few-shot segmentation method based on generalization features and a region proposal loss function. For the input image to be segmented, first, various data augmentations and the training of a deformable convolutional network are performed to make the model have stronger generalization ability. Then, channel attention and spatial attention modules are used to extract more advanced deep features. Finally, the edge features of the image are extracted and region proposals are introduced into the model, fully considering the edges and accurate positioning of the image. Finally, a few-shot segmentation mechanism is introduced to fine-tune the model under the premise of low data volume to achieve accurate image segmentation results. So as to be able to effectively perform image segmentation with fewer sample trainings.
[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0006] A few-shot segmentation method based on generalization features and a region proposal loss function, characterized by comprising the following steps:
[0007] Step S1: Split the general training set with a small sample, and train the feature extraction network using a data augmentation method to improve feature generalization;
[0008] Step S2: Use the trained feature extraction network as the feature extraction module, introduce the attention mechanism, and deepen the foreground part in the features;
[0009] Step S3: Use the obtained foreground part features to extract edge features from the image, and obtain the image foreground features and the edge features of the image;
[0010] Step S4: Use the region proposal loss function to further locate the position of the object, and perform calculations using the location and edge features to complete small sample segmentation.
[0011] Further, step S1 specifically includes the following steps:
[0012] Step S11: Obtain the publicly available small sample segmentation training set from the network and obtain the relevant annotations of the training data; divide the data set into a base class and a new class according to the category. Among them, the base class is a data set containing a large amount of data and a large number of categories, and the new class is a category that has never appeared in the base class data set, and the number of new classes is less than that of the base class; the divided new class data set is used for model fine-tuning in step S4;
[0013] Step S12: Perform data augmentation on the image to be segmented. First, to improve the generalization performance, divide the picture into ten rectangles with the same length and width in the horizontal and vertical directions, and remove 30% of each of the ten regions in the horizontal and vertical directions according to the probability, and fill the removed regions with black pixels, and introduce histogram equalization to make the image color features of the image to be segmented more obvious. Finally, offset the region of the segmented object in the annotation by 5% around, and obtain the image to be segmented after image enhancement;
[0014] Step S13: Initialize the network weights and parameters for the segmentation images in the small sample segmentation training set using the pre-trained ResNet101 feature extraction network to accelerate the model convergence speed;
[0015] Step S14: Use the 101-layer ResNet as the network model for extracting small sample image generalization features, and replace the third and fourth layer residual networks with deformable convolutional networks: First, increase the traditional 3×3 convolutional kernel by offsetting {Δp n |n = 1,..., N}, where N = |R|, R represents the set of real numbers, and use the increased convolutional kernel to upsample on the input feature map x feature and sum the sampled values weighted by the learnable parameter w. For each position p output on the output feature map y 0 , obtain the output feature map y outputThe calculation formula is as follows:
[0016]
[0017] R p = {(-1, -1), (-1, 0), …, (0, 1), (1, 1)}
[0018] where R p defines the size of the receptive field and the range of dilation, p n enumerates the specific positions listed in R p , n is the given number of offsets in all directions, and N is the number of offsets required for the deformable convolutional kernel;
[0019] Step S15: After introducing the deformable convolutional network, the sampling of the input feature is at the irregular and offset position p n + Δp n The output feature map y output is calculated by bilinear interpolation to obtain specific values, and its calculation formula is as follows:
[0020]
[0021] G(q, p) = g(q x , p x ) · g(q y , p y )
[0022] where p represents any position, i.e., p = p 0 + p n + Δp n , q enumerates all integer positions in the feature map x feature , G is the kernel of bilinear interpolation, which is calculated by dividing it into two one-dimensional kernels, g is the calculation function of one-dimensional linear interpolation, q x , p x corresponds to the abscissa value of any position p in the feature map and all integer positions q in the feature map x feature , q y , p y corresponds to the ordinate value of any position p in the feature map and all integer positions q in the feature map x feature .
[0023] Furthermore, Step S2 includes the following steps:
[0024] Step S21: Use the trained ResNet101 network in Step S1 as the feature extraction model to extract the depth feature F of each segmented image on the same training dataset 1 ∈ R C×H×W, where \(R\) represents the set of all real numbers, and \(C\), \(H\), and \(W\) represent the number of channels of the feature, as well as the height and width of the feature matrix respectively; the resulting feature map of the final output of the deformable convolutional network is used as the input to the attention mechanism module.
[0025] Step S22: To find the correlation between the features and the segmentation targets, a channel attention module is used to learn the weights of the generated features: First, the channel attention sub-module preserves information in three dimensions using three-dimensional permutations, uses a three-layer multi-layer perceptron to amplify the cross-dimensional channel-spatial dependence, and adopts Relu as the activation function. The image feature after passing through the channel attention sub-module is \(F\) 2 , \(F\) 2 is calculated as follows:
[0026]
[0027] where \(M\) c represents the three-layer perceptron of channel attention and is an encoder-decoder structure; represents element-wise multiplication operation;
[0028] Step S23: To amplify the global dimension interaction features while reducing information loss, a spatial attention sub-module is introduced. Two grouped convolutional layers with Channel Shuffle are used for spatial information fusion, the pooling operation is directly removed, Relu is adopted as the activation function, and a batch normalization layer is inserted between the fully connected layer and the activation function for the normalization operation of the input; the image feature after passing through the spatial attention sub-module is \(F\) 3 , \(F\) 3 is calculated as follows:
[0029]
[0030] where \(M\) s represents the two-layer perceptron of spatial attention and is an encoder-decoder structure.
[0031] Furthermore, Step S3 specifically includes the following steps:
[0032] Step S31: Use a linear projection to predict the query coefficients of each segmentation target:
[0033] \(Q'=\text{sigmoid}(\text{linear}(d,f)(F 3 ))
[0034] Among them, sigmoid(·) represents the activation function in the neural network, which can map the query coefficient to the interval of (0, 1). Both d and f represent the numerical values of dimensions and are integers. linear(d, f) represents the linear projection from d dimensions to f dimensions, and Q′ is a query coefficient for each segmentation target.
[0035] Step S32: To predict the edge map of each target to be segmented, use the query coefficient Q′ of the segmentation target j as the filter weight to perform 1×1 convolution on the feature map F 3 ; which is equivalent to applying a batch matrix multiplication between Q′ and F 3 :
[0036]
[0037] Among them, j is the index of each segmentation target, and O j is the edge map of each segmentation target; all segmentation targets are multiplied by the same set of feature maps.
[0038] Furthermore, step S4 specifically includes the following steps:
[0039] Step S41: Use the region proposal loss function to calculate the difference between the predicted position of the object to be segmented and the actual position and obtain the loss value L TIoU (F 3 ), and its calculation formula is as follows:
[0040]
[0041]
[0042] Among them, IoU represents the intersection over union of the predicted localization and the actual localization of the segmentation target. A and B respectively represent the predicted box and the ground truth box. ρ represents the Euclidean distance between the two center points. b pre and b gt respectively represent the center points of the predicted box and the ground truth box. c represents the diagonal distance of the smallest closed region that can contain both the predicted box and the ground truth box. w pre , h pre are respectively the width and height of the predicted box, and w gt , h gt are respectively the width and height of the ground truth box. represents the sum of the widths and heights of the smallest boxes in the intersection of the predicted box and the ground truth box;
[0043] Step S42: Test the trained model on the validation set to obtain the final segmentation accuracy; on the basis of the trained model, introduce the mechanism of small samples, freeze all the module parameters before step S3, and perform fine-tuning including steps S3 and S41 on the new class dataset to obtain a training model adapted to the segmentation of the new class dataset.
[0044] Compared with the prior art, the present invention and its preferred embodiments have the following beneficial effects:
[0045] 1. It can converge quickly with fewer samples, obtain better segmentation performance, and improve the generalization ability of the model.
[0046] 2. It can start from the aspect of feature extraction and improve the error correction ability of the model for images through delicate operations of data augmentation.
[0047] 3. Since the shape and appearance of the objects to be concerned about in image segmentation have a great influence on model training, deformable convolution is introduced. After having deeper features in the last layer of feature extraction, through the offset calculation of deformable convolution, the relationship between the surrounding of the image and the segmented object can be considered, and the position of the segmented object can be better determined to improve the classification and segmentation performance of the segmented object.
[0048] 4. Aiming at the fact that in the traditional Smooth L1 loss function, the regression function of the target to be segmented is difficult to accurately predict, a loss function based on region proposals is proposed. The difference between the predicted box and the ground truth box is calculated, and the height and width of the anchor box are introduced as considerations in loss calculation, making the prediction of the anchor box more accurate. For the small sample segmentation problem, it can well solve the problem of data overfitting and improve the generalization of the model. Description of the Drawings
[0049] The present invention will be further described in detail below with reference to the drawings and specific embodiments:
[0050] Figure 1 is the flowchart of the method implementation of the embodiment of the present invention. Detailed Embodiments
[0051] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below and described in detail as follows:
[0052] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0053] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0054] As Figure 1 shown, an embodiment of the present invention provides a few-shot segmentation method based on generalization features and a regional proposal loss function, including the following steps:
[0055] Step S1: Train a feature extraction network by using a few-shot segmentation general training set and a data augmentation method for improving feature generalization. Specifically, it includes the following steps:
[0056] Step S11: Obtain a publicly available few-shot segmentation training set from the network and obtain the relevant annotations of the training data. Divide the data set into a base class and a new class according to the categories. The base class is a data set containing a large amount of data and a large number of categories, while the new class is a category that has never appeared in the base class data set, and the number of new classes is much smaller than that of the base class. The divided new class data set is used in the model fine-tuning in step S4;
[0057] Step S12: Perform data augmentation on the image to be segmented. First, to improve the generalization performance, divide the picture into ten rectangles with the same length and width in the horizontal and vertical directions, and remove 30% of each of the ten regions in the horizontal and vertical directions according to the probability, and fill the removed regions with black pixels, and introduce histogram equalization to make the image color features of the image to be segmented more obvious. Finally, offset the region of the segmented object in the annotation by 5% around, and obtain the image to be segmented after image enhancement;
[0058] Step S13: Use a pre-trained ResNet101 feature extraction network to initialize the network weights and parameters for the segmentation images in the few-shot segmentation training set to accelerate the model convergence speed;
[0059] Step S14: Use a 101-layer ResNet as the network model for extracting few-shot image generalization features, and replace the residual networks of the third and fourth layers with deformable convolutional networks. First, increase the traditional 3×3 convolutional kernel by the offset {Δp n |n = 1,..., N}, where N = |R|, and R represents the set of real numbers. Use the increased convolutional kernel to upsample the input feature map x feature and sum the sampled values weighted by the learnable parameter w. For each position p output on the output feature map y 0 , the output feature map y can be obtainedoutput The calculation formula is as follows:
[0060]
[0061] R p = {(-1, -1), (-1, 0), …, (0, 1), (1, 1)}
[0062] where R p defines the size of the receptive field and the range of dilation, p n enumerates the specific positions listed in R p , n is the given number of offsets in all directions, and N is the number of offsets required for the deformable convolution kernel;
[0063] Step S15: After introducing the deformable convolution network, the sampling of the input features is now at the irregular and offset positions p n + Δp n and the output feature map y output can be calculated by bilinear interpolation to obtain specific values, and its calculation formula is as follows:
[0064]
[0065] G(q, p) = g(q x, p x ) · g(q y , p y )
[0066] where p represents any position, i.e., p = p 0 + p n + Δp n , q enumerates all integer positions in the feature map x feature , G is the kernel of bilinear interpolation, which is divided into two one-dimensional kernels for calculation, g is the calculation function of one-dimensional linear interpolation, q x , p x corresponds to the abscissa value of any position p in the feature map and all integer positions q in the feature map x feature , q y , p y corresponds to the ordinate value of any position p in the feature map and all integer positions q in the feature map x feature .
[0067] Step S2: Use the trained feature extraction network as the feature extraction module, introduce the attention mechanism while having the basic generalization performance, and deepen the foreground part in the features. Specifically, it includes the following steps:
[0068] Step S21: Use the trained ResNet101 network in Step S1 as the feature extraction model to extract the depth feature F of each segmented image on the same training dataset 1 ∈R C×H×W , where R represents the set of all real numbers, and C, H, and W represent the number of channels of the feature, and the height and width of the feature matrix, respectively. Output the result of the feature map of the final output of the deformable convolutional network as the input to the attention mechanism module;
[0069] Step S22: To find the correlation between the feature and the segmentation target, a channel attention module can be used to learn and generate the weight of the feature. First, the channel attention sub-module preserves information in three dimensions using three-dimensional permutation, uses a three-layer multi-layer perceptron to amplify the cross-dimensional channel-space dependence, and uses Relu as the activation function. The image feature after passing through the channel attention sub-module is F 2 , F 2 The calculation method of is as follows:
[0070]
[0071] Among them, M c represents the three-layer perceptron of channel attention, which is an encoder-decoder structure. represents element-wise multiplication operation;
[0072] Step S23: To amplify the global dimension interaction feature while reducing information loss, a spatial attention sub-module is introduced, and two grouped convolutional layers with Channel Shuffle are used for spatial information fusion. Since the max-pooling operation reduces the use of information, the pooling operation is directly removed, Relu is used as the activation function, and a batch normalization layer is inserted between the fully connected layer and the activation function for input normalization operation. The image feature after passing through the spatial attention sub-module is F 3 , F 3 The calculation method of is as follows:
[0073]
[0074] Among them, M s represents the two-layer perceptron of spatial attention, which is an encoder-decoder structure.
[0075] Step S3: Use the obtained foreground part features to extract the edge features of the image, and finally obtain the foreground features of the image and the edge features of the image. Specifically, it includes the following steps:
[0076] Step S31: The depth target feature F output by the attention module 3It can well distinguish foreground and background features. However, in order to more precisely obtain the edges of the segmented objects, it is still necessary to predict the object edges. First, a simple linear projection is used to predict the query coefficients for each segmented object:
[0077] Q′ = sigmoid(linear(d, f)(F 3 ))
[0078] where sigmoid(·) represents the activation function in the neural network, which can map the query coefficients to the interval (0, 1). Both d and f represent the numerical values of dimensions and are integers. linear(d, f) represents the linear projection from dimension d to dimension f, and Q′ is a query coefficient for each segmented object;
[0079] Step S32: To predict the edge map for each object to be segmented, use the query coefficient Q′j of the segmented object as the filter weight to perform 1×1 convolution on the feature map F 3 This is equivalent to applying a batch matrix multiplication between Q′ and F 3 :
[0080]
[0081] where j is the index of each segmented object, and O j is the edge map of each segmented object. All segmented objects are multiplied by the same set of feature maps.
[0082] Step S4: Use the region proposal loss function to further localize the position of the object, and calculate using the localization and edge features to complete the final few-shot segmentation. Specifically, it includes the following steps:
[0083] Step S41: After obtaining the deep features F 3 of the image to be segmented and the edge map O j of the segmented object, in order to better and accurately localize the segmented object, use the region proposal loss function to calculate the difference between the predicted position and the actual position of the object to be segmented and obtain the loss value L TIoU (F 3 ), and its calculation formula is as follows:
[0084]
[0085]
[0086] where IoU represents the intersection over union of the predicted localization and the actual localization of the segmented object, A and B represent the predicted bounding box and the ground truth bounding box respectively, ρ represents the Euclidean distance between the two center points, b pre and bgt respectively represent the center points of the predicted bounding box and the ground truth bounding box, c represents the diagonal distance of the smallest closed region that can contain both the predicted bounding box and the ground truth bounding box, w pre , h pre are the width and height of the predicted bounding box respectively, w gt , h gt are the width and height of the ground truth bounding box respectively, represents the sum of the width and height of the smallest bounding box of the intersection in the predicted bounding box and the ground truth bounding box;
[0087] Step S42: Test the trained model on the validation set to obtain the final segmentation accuracy. Based on the trained model, introduce a small sample mechanism, freeze all the module parameters before step S3, and fine-tune on the training set that the model has never seen, namely the new class dataset. The part for fine-tuning is the operations after step S3 and step S3 to obtain a training model adapted to the segmentation of the new class dataset, and the fine-tuned model still has good segmentation performance.
[0088] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0089] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0090] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more flows and / or blocks Figure 1The functions specified in one or more boxes.
[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one or more processes and / or boxes Figure 1 One process or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes.
[0092] As described above, it is only the preferred embodiment of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
[0093] This patent is not limited to the above best implementation manner. Anyone inspired by this patent can obtain various other forms of small sample segmentation methods based on the generalization feature and region proposal loss function. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the coverage scope of this patent.
Claims
1. A few-shot segmentation method based on generalized features and region proposal loss function, characterized in that, it includes the following steps: Step S1: Train the feature extraction network by a few-shot segmentation general training set and adopt a data augmentation method to improve feature generalization; Step S2: Use the trained feature extraction network as the feature extraction module, introduce the attention mechanism, and deepen the foreground part in the features; Step S3: Use the obtained foreground part features to extract edge features of the image, and obtain the foreground features of the image and the edge features of the image; Step S4: Use the region proposal loss function to further locate the position of the object, calculate using the location and edge features, and complete the few-shot segmentation.
2. The few-shot segmentation method based on generalized features and region proposal loss function according to claim 1, characterized in that, Step S1 specifically includes the following steps: Step S11: Obtain the publicly available few-shot segmentation training set from the network and obtain the relevant annotations of the training data; divide the data set into base classes and new classes according to categories, where the base class is a data set containing a large amount of data and a large number of categories, and the new class is a category that has never appeared in the base class data set, and the number of new classes is less than that of the base class; the divided new class data set is used for model fine-tuning in Step S4; Step S12: Perform data augmentation on the image to be segmented. First, to improve the generalization performance, divide the picture into ten rectangles with the same length and width in both the horizontal and vertical directions, and remove 30% of each of the ten regions in the horizontal and vertical directions according to probability, and fill the removed regions with black pixels, and introduce histogram equalization to make the image color features of the image to be segmented more obvious. Finally, offset the region of the segmented object in the annotation by 5% around, and obtain the image to be segmented after image enhancement; Step S13: Initialize the network weights and parameters of the segmentation images in the few-shot segmentation training set by using the pre-trained ResNet101 feature extraction network to accelerate the model convergence speed; Step S14: Use a 101-layer ResNet as the network model for extracting the generalization features of few-shot images, and replace the residual networks of the third and fourth layers with deformable convolutional networks: First, increase the traditional 3×3 convolutional kernel by the offset {Δp n |n = 1,…,N}, where N = |R| and R represents the set of real numbers, and use the increased convolutional kernel to perform upsampling on the input feature map x feature , and sum the sampled values weighted by the learnable parameter w. For each position p output on the output feature map y 0 , the calculation formula for the output feature map y output is as follows: R p ={(-1,-1),(-1,0),…,(0,1),(1,1)} where R p defines the size of the receptive field and the range of dilation, p n enumerates the specific positions listed in R p the number of given offsets around, and N is the number of offsets required for the deformable convolution kernel; Step S15: After introducing the deformable convolutional network, the sampling of the input features is performed at the irregular and offset positions p n +Δp n and the output feature map y output is calculated by bilinear interpolation to obtain specific values, and its calculation formula is as follows: G(q,p) = g(q x ,p x )·g(q y ,p y ) where p represents any position, i.e., p = p 0 + p n + Δp n , q enumerates all integer positions in the feature map x feature . G is the kernel of bilinear interpolation, which is calculated by dividing it into two one-dimensional kernels. g is the calculation function of one-dimensional linear interpolation, q x , p x corresponds to the abscissa value of any position p in the feature map and all integer positions q in the feature map x feature . q y , p y corresponds to the ordinate value of any position p in the feature map and all integer positions q in the feature map x feature .
3. The few-shot segmentation method based on generalized features and region proposal loss function according to claim 2, characterized in that, Step S2 includes the following steps: Step S21: Use the trained ResNet101 network in Step S1 as a feature extraction model to extract the depth feature F of each segmented image on the same training dataset 1 ∈R C×H×W , where R represents the set of all real numbers, and C, H, and W represent the number of channels of the feature, the height, and the width of the feature matrix, respectively; Output the feature map result of the final output of the deformable convolutional network as the input of the attention mechanism module; Step S22: To find the correlation between features and segmentation targets, a channel attention module is used to learn the weights of the generated features. First, the channel attention sub-module preserves information in three dimensions using a three-dimensional arrangement, uses a three-layer multi-layer perceptron to amplify the cross-dimensional channel-spatial dependence, and adopts Relu as the activation function. The image feature passing through the channel attention sub-module is F 2 , F 2 is calculated as follows: Among them, M c represents a three-layer perceptron for channel attention and has an encoder-decoder structure; represents element-wise multiplication operation; Step S23: To amplify the global dimension interaction features while reducing information loss, a spatial attention sub-module is introduced. Two grouped convolutional layers with Channel Shuffle are used for spatial information fusion, the pooling operation is directly removed, Relu is used as the activation function, and a batch normalization layer is inserted between the fully connected layer and the activation function for input normalization; the image features after passing through the spatial attention sub-module are F 3 , F 3 is calculated as follows: Among them, M s represents a two-layer perceptron for spatial attention and has an encoder-decoder structure.
4. The few-shot segmentation method based on generalized features and region proposal loss function according to claim 3, characterized in that, Step S3 specifically includes the following steps: Step S31: Use a linear projection to predict the query coefficient of each segmentation target: Q' = sigmoid(linear(d,f)(F 3 )) where sigmoid(·) represents the activation function in the neural network, which can map the query coefficient to the interval of (0,1), d and f both represent the numerical values of dimensions and are both integers; linear(d,f) represents the linear projection from d dimensions to f dimensions, and Q' is a query coefficient of each segmentation target; Step S32: To predict the edge map of each target to be segmented, use the query coefficient Q' of the segmentation target j as the filter weights and perform a 1×1 convolution on the feature map F 3 ; which is equivalent to applying a batch matrix multiplication between Q' and F 3 : where j is the index of each segmented object, and O j is the edge map of each segmented object; all segmented objects are multiplied by the same set of feature maps.
5. The few-shot segmentation method based on generalized features and region proposal loss function according to claim 4, characterized in that, Step S4 specifically includes the following steps: Step S41: Calculate the difference between the predicted position of the object with segmentation and the actual position using the region proposal loss function to obtain a loss value L TIoU (F 3 ), and its calculation formula is as follows: Among them, IoU represents the intersection over union of the predicted localization and the actual localization of the segmentation target. A and B represent the predicted bounding box and the ground truth bounding box respectively. ρ represents the Euclidean distance between the two center points, and b pre and b gt represent the center points of the predicted bounding box and the ground truth bounding box respectively. c represents the diagonal distance of the smallest closed region that can contain both the predicted bounding box and the ground truth bounding box. w pre , h pre are the width and height of the predicted bounding box respectively. w gt , h gt are the width and height of the ground truth bounding box respectively. represents the sum of the width and height of the smallest bounding box of the intersection between the predicted bounding box and the ground truth bounding box; Step S42: Test the trained model on the validation set to obtain the final segmentation accuracy; on the basis of the trained model, introduce a few-shot mechanism, freeze the module parameters before step S3, and perform fine-tuning including steps S3 and S41 on the new class dataset to obtain a training model adapted to the segmentation of the new class dataset.
Citation Information
Patent Citations
MRI (Magnetic Resonance Imaging) segmentation method for integrating attention mechanism aiming at brain lesion
CN114332462A
Small sample target detection method based on double-branch region suggestion network
CN114743045A