SAM-based butterfly ecological image segmentation method and system
By adding a dual-channel convolution module and feature fusion module to the SAM model, the problems of insufficient accuracy and high training cost in butterfly ecological image segmentation are solved, and higher segmentation accuracy and lower training cost are achieved.
Patent Information
- Application Number
- CN202510093583.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The prior art has problems in the segmentation of butterfly ecological image in the butterfly ecological image, such as insufficient accuracy, broken boundaries, and confusion of butterfly pixels and background pixels. The SAM model has a huge amount of parameters and high training costs.
Based on the SAM model, a dual-channel convolution module and a feature fusion module are added, and the features generated by the image encoder are further extracted through the dual-channel convolution module to obtain global and local features, and the global and mask features are fused through the feature fusion module to enrich the details of the mask features.
The accuracy of butterfly ecological image segmentation has been improved and the parameters and cost of model training have been reduced. The IoU index has increased from 89.82% to 91.78%, and the MIoU index has increased from 94.76% to 95.78%.
Smart Images

Figure CN120013964A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image segmentation, and in particular relates to a butterfly ecological image segmentation method and system based on SAM. Background Art
[0002] In butterfly ecological images, the mimicry of butterflies makes them very similar to the background. Therefore, the automatic recognition accuracy of butterfly ecological images with background is much lower than that of butterfly ecological images without most of the background. The existing image segmentation network has problems such as insufficient accuracy, broken boundaries, and confusion between butterfly pixels and background pixels in butterfly ecological images. Therefore, a special segmentation model for butterfly ecological images is needed.
[0003] Kirillov, Alexander et al. proposed the SAM model in the paper "Segment Anything". The model consists of an image encoder, an image decoder, and a prompt encoder. The image is input into the image encoder, the point, box, and text prompts are input into the prompt encoder, and then the features generated by the image encoder and the prompt encoder are input into the image decoder to obtain the butterfly mask. However, SAM still lacks the accuracy of butterfly ecological image segmentation, and the number of parameters is huge, making it difficult to train. Summary of the invention
[0004] In order to solve the problems existing in the prior art, the present invention provides a butterfly ecological image segmentation method based on SAM, aiming to solve the problems of unsatisfactory segmentation result accuracy and high training cost in butterfly ecological image segmentation of SAM, by adding a two-way convolution module, further extracting features generated by the image encoder, obtaining global features and local features of the image, and by adding a feature fusion module, fusion of the global features of the image and the mask features generated by the image decoder, enriching the details of the mask features, and improving the segmentation accuracy of the model for butterfly ecological images. By freezing the parameters of the image encoder, the prompt encoder, and the mask decoder, only the parameters of the two-way convolution module and the feature fusion module with a small number of parameters are trained.
[0005] In order to achieve the above object, in a first aspect, the present invention provides a butterfly ecological image segmentation method based on SAM, comprising the following steps: The resized butterfly ecological image is annotated using a mask to obtain a binary butterfly ecological image; The original butterfly ecological image is input into the butterfly segmentation network; the butterfly segmentation network includes an image encoder, a prompt encoder, an image decoder, a two-way convolution module, and a feature fusion module; the image encoder is a pre-trained image encoder of the SAM model, which is composed of 12 Transformer modules connected in sequence, and is used to encode the butterfly ecological image, input the butterfly ecological image, and input the image feature vector output by the last Transformer module into the image decoder; the prompt encoder is a pre-trained prompt encoder of the SAM model, which is used to encode the point, text, and anchor box prompts into a prompt vector; the image decoder is a pre-trained image decoder of the SAM model, which is used to decode the butterfly ecological image feature vector generated by the image encoder and the prompt vector generated by the prompt encoder to generate a mask feature vector; the two-way convolution module is used to further extract features from the feature coding vector generated by the 12 Transformer modules of the image encoder, including two pathways, namely the global pathway and the local pathway. The global pathway uses the image feature vectors generated by the 3rd, 6th, 9th, and 12th Transformer modules of the image encoder to generate an image global feature vector, and the local pathway uses the image feature vectors generated by the remaining Transformer modules to generate an image local feature vector; The global feature vector, local feature vector and mask feature vector of the image are fused to obtain the final butterfly mask image.
[0006] Furthermore, the resized butterfly ecological image is annotated using a mask, and the binary butterfly ecological image obtained includes: The butterfly ecological image is adjusted to an image with a height of H and a width of W. C represents the image channel. The number of color RGB image channels is 3. The butterfly ecological image is recorded as , use mask method to mark, use labelimg software to make several points on the butterfly boundary, connect the several points with line segments to form a closed image, the closed image takes value 1 inside and value 0 outside, and obtains a png image composed of 0 and 1, where the area with value 1 is the butterfly area.
[0007] Furthermore, when training the butterfly segmentation network, the parameters of the image encoder, the hint encoder, and the image decoder are frozen, and only the parameters of the two-way convolution module and the feature fusion module are trained; the loss function used for training the parameters of the two-way convolution module and the feature fusion module is:
[0008] in Represents the binary cross entropy loss function BCE (Binary CrossEntropy Loss), which is determined by the following formula;
[0009] in, is a specific pixel point, Pixel The probability of predicting the butterfly class, for Corresponding to the true label, is the total number of pixels in an image; represents the Dice loss function, which is determined by the following formula;
[0010] in, is a specific pixel point, Pixel The probability of predicting the butterfly class, for Corresponding to the true label, is the total number of pixels in an image.
[0011] Furthermore, the global path is connected in sequence by 4 basic convolution modules, the image feature vector generated by the 3rd Transformer module of the image encoder is input into the 1st basic convolution module of the global path, the image feature vector output by the 1st basic convolution module is added to the image feature vector generated by the 6th Transformer module of the image encoder, and then input into the 2nd basic convolution module of the global path, the image feature vector output by the 2nd basic convolution module is added to the image feature vector generated by the 9th Transformer module of the image encoder, and then input into the 3rd basic convolution module of the global path, the image feature vector output by the 3rd basic convolution module is added to the image feature vector generated by the 12th Transformer module of the image encoder, and then input into the 4th basic convolution module of the global path, and the image feature vector generated by it is the image global feature vector.
[0012] Furthermore, the local pathway is connected in sequence by 8 basic convolution modules, the image feature vector generated by the first Transformer module of the image encoder is input into the first basic convolution module of the local pathway, the image feature vector output by the first basic convolution module is added to the image feature vector generated by the second Transformer module of the image encoder, and then input into the second basic convolution module of the local pathway, the image feature vector output by the second basic convolution module is added to the image feature vector generated by the fourth Transformer module of the image encoder, and then input into the third basic convolution module of the local pathway, the image feature vector output by the third basic convolution module is added to the image feature vector generated by the fifth Transformer module of the image encoder, and then input into the fourth basic convolution module of the local pathway, and the fourth basic convolution module is added. After adding the image feature vector output by the block to the image feature vector generated by the 7th Transformer module of the image encoder, the image feature vector is input into the 5th basic convolution module of the local pathway. After adding the image feature vector output by the 5th basic convolution module to the image feature vector generated by the 8th Transformer module of the image encoder, the image feature vector is input into the 6th basic convolution module of the local pathway. After adding the image feature vector output by the 6th basic convolution module to the image feature vector generated by the 10th Transformer module of the image encoder, the image feature vector is input into the 7th basic convolution module of the local pathway. After adding the image feature vector output by the 7th basic convolution module to the image feature vector generated by the 11th Transformer module of the image encoder, the image feature vector is input into the 8th basic convolution module of the local pathway. The generated image feature vector is the image local feature vector.
[0013] Furthermore, the basic convolution module is used to extract features from the image feature vector generated by the image encoder Transformer module, and is composed of a convolution layer with a size of 3*3 and a step size of 1, a normalization layer, a relu activation function layer, a convolution layer with a size of 3*3 and a step size of 1, a normalization layer, and a relu activation function layer.
[0014] Furthermore, the feature fusion module is composed of a learnable 1*256-dimensional vector and a three-layer perceptron. The 1*256-dimensional vector is processed by the three-layer perceptron to become a 1*32-dimensional vector. After adding the global feature vector, the local feature vector and the mask vector generated by the two-way convolution module and the mask decoder, the 1*32-dimensional vector is used to perform matrix multiplication therewith to obtain the final segmented butterfly mask image.
[0015] In a second aspect, the present invention also provides a butterfly ecological image segmentation system based on SAM, including an image preprocessing module, a feature extraction module, and a feature fusion module; The image preprocessing module is used to mark the resized butterfly ecological image using a mask to obtain a binary butterfly ecological image; The feature extraction module is used to input the original butterfly ecological image into the butterfly segmentation network; the butterfly segmentation network includes an image encoder, a prompt encoder, an image decoder, a two-way convolution module, and a feature fusion module; the image encoder is a pre-trained image encoder of the SAM model, which is composed of 12 Transformer modules connected in sequence, and is used to encode the butterfly ecological image, input the butterfly ecological image, and input the image feature vector output by the last Transformer module into the image decoder; the prompt encoder is a pre-trained prompt encoder of the SAM model, which is used to encode the point, text, and anchor box prompts into a prompt vector; the image decoder is a pre-trained image decoder of the SAM model, which is used to decode the butterfly ecological image feature vector generated by the image encoder and the prompt vector generated by the prompt encoder to generate a mask feature vector; the two-way convolution module is used to further extract features from the feature coding vector generated by the 12 Transformer modules of the image encoder, including two pathways, namely the global pathway and the local pathway. The global pathway uses the image feature vectors generated by the 3rd, 6th, 9th, and 12th Transformer modules of the image encoder to generate an image global feature vector, and the local pathway uses the image feature vectors generated by the remaining Transformer modules to generate an image local feature vector; The feature fusion module is used to fuse the image global feature vector, local feature vector and mask feature vector to obtain the final butterfly mask image.
[0016] In a third aspect, the present invention can also provide a computer device, including a processor and a memory, the memory is used to store a computer executable program, the processor reads the computer executable program from the memory and executes it, and when the processor executes the computer executable program, the SAM-based butterfly ecological image segmentation method described in the present invention can be implemented.
[0017] A computer-readable storage medium is also provided, in which a computer program is stored. When the computer program is executed by a processor, the butterfly ecological image segmentation method based on SAM described in the present invention can be implemented.
[0018] Compared with the prior art, the present invention has at least the following beneficial effects: SAM needs to train the image encoder, prompt encoder and mask decoder, and the required training parameter amount is 1191M. The present invention only needs to train the dual-path convolution module and the feature fusion module, and the required training parameter amount is about 14.7M, which is much lower than the training parameter amount required for SAM. The training can be completed on a GPU. Therefore, the training cost of the present invention is much lower than retraining the SAM model.
[0019] The image coding features generated by each Transformer module in the image encoder are input into the dual-path convolution module for feature extraction again, the information generated by each Transformer module is fully utilized, more knowledge of image details is learned, and image global features and image local features are generated. The image global features, image local knowledge and image mask features generated by the mask decoder are input into the feature fusion module and fused in an addition manner, which enriches the details of the image mask features, and converts the image mask features into the final segmentation mask output through a learnable vector, thereby improving the segmentation accuracy of butterfly ecological images. Compared with the SAM model, the IoU index of butterfly ecological image segmentation of the present invention is improved from 89.82% to 91.78%, and the MIoU index is improved from 94.76% to 95.78%. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 The present invention provides a flowchart of an embodiment.
[0021] Figure 2 The overall flow chart provided by the present invention.
[0022] Figure 3 This is a flow chart of the dual-path convolution module provided by the present invention.
[0023] Figure 4 It is a global pathway flow chart provided by the present invention.
[0024] Figure 5 It is a local pathway flow chart provided by the present invention.
[0025] Figure 6 This is a basic convolution module flow chart provided by the present invention.
[0026] Figure 7 It is a flow chart of the feature fusion module provided by the present invention.
[0027] Figure 8 It is a schematic diagram of the image segmentation result provided by the present invention. DETAILED DESCRIPTION
[0028] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0029] In the description of the present invention, it should be understood that the terms “include” and “comprises” indicate the presence of described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.
[0030] refer to Figure 1 The present invention provides a butterfly ecological image segmentation method based on SAM, comprising the following steps: S1, construct a butterfly segmentation dataset, adjust the butterfly pictures to images with a height of H and a width of W pixels, marked as , C represents the three channels of the image. Labelimg software is used to make several points on the butterfly boundary, and these points are connected by line segments to form a closed image. The value inside the closed image is 1, and the value outside is 0. A png image composed of 0 and 1 is obtained, where the area with a value of 1 is the butterfly area. The data set is divided into two parts at a ratio of 1:1, one as a training set and the other as a test set; S2, construct a butterfly segmentation network. The butterfly segmentation network includes a dual-path convolution module and a feature fusion module based on the image encoder, prompt encoder and image decoder originally possessed by SAM; the image encoder is connected by 12 Transformer modules in sequence, the image is input to the image encoder, each Transformer module can obtain an image feature vector, the image feature vector generated by the 12th Transformer module, the prompt encoder generates a prompt vector, the image feature vector and the prompt vector are input to the mask decoder together, the mask decoder generates a mask feature vector, and the 12 image feature vectors generated by the 12 Transformer modules are input to the dual-path convolution module, the dual-path convolution module includes two pathways, a global pathway and a local pathway, the global pathway consists of 4 basic convolutions The convolution modules are connected in sequence, the image feature vector generated by the third Transformer module of the image encoder is input into the first basic convolution module of the global path, the image feature vector output by the first basic convolution module is added to the image feature vector generated by the sixth Transformer module of the image encoder, and then input into the second basic convolution module of the global path, the image feature vector output by the second basic convolution module is added to the image feature vector generated by the ninth Transformer module of the image encoder, and then input into the third basic convolution module of the global path, the image feature vector output by the third basic convolution module is added to the image feature vector generated by the twelfth Transformer module of the image encoder, and then input into the fourth basic convolution module of the global path, and the image feature vector generated by it is the image global feature vector;The local pathway is connected in sequence by 8 basic convolution modules. The image feature vector generated by the first Transformer module of the image encoder is input into the first basic convolution module of the local pathway. The image feature vector output by the first basic convolution module is added to the image feature vector generated by the second Transformer module of the image encoder, and then input into the second basic convolution module of the local pathway. The image feature vector output by the second basic convolution module is added to the image feature vector generated by the fourth Transformer module of the image encoder, and then input into the third basic convolution module of the local pathway. The image feature vector output by the third basic convolution module is added to the image feature vector generated by the fifth Transformer module of the image encoder, and then input into the fourth basic convolution module of the local pathway. The image feature vector output by the fourth basic convolution module is added to the image feature vector generated by the seventh Transformer module of the image encoder, and then input into the fifth basic convolution module of the local pathway. The image feature vector output by the basic convolution module is added to the image feature vector generated by the 8th Transformer module of the image encoder, and then input into the 6th basic convolution module of the local path. The image feature vector output by the 6th basic convolution module is added to the image feature vector generated by the 10th Transformer module of the image encoder, and then input into the 7th basic convolution module of the local path. The image feature vector output by the 7th basic convolution module is added to the image feature vector generated by the 11th Transformer module of the image encoder, and then input into the 8th basic convolution module of the local path. The image feature vector generated is the local feature vector of the image. The global feature vector, local feature vector, and mask feature vector of the image are input into the feature fusion module. After the feature fusion module adds the global feature vector, local feature vector, and mask feature vector of the image, the learnable 1*256-dimensional vector is first reduced to 1*32-dimensional, and then the 1*32-dimensional vector is multiplied with the fused feature to obtain the final butterfly mask image output; reference; Figure 3 , Figure 4 , Figure 5 and Figure 6 .
[0031] S3, construct the loss function according to the following formula ;
[0032] in Represents the binary cross entropy loss function BCE (Binary CrossEntropy Loss), which is determined by the following formula;
[0033] in, is a specific pixel point, Pixel The probability of predicting the butterfly class, for Corresponding to the true label, is the total number of pixels in an image; represents the Dice loss function, which is determined by the following formula;
[0034] in, is a specific pixel point, Pixel The probability of predicting the butterfly class, for Corresponding to the true label, is the total number of pixels in an image; S4, input the training set into the butterfly detection network for training, freeze the parameters of the image encoder, prompt encoder, and image decoder, and only train the parameters of the two-way convolution module and the feature fusion module. The server graphics card for training is RTX3090, the initial learning rate is 0.001, the number of training rounds is 20, the training batch size is 4, and the training is carried out until the loss function converges; S5, save the converged weight file.
[0035] S6, the butterfly ecological image network loads the saved weight file, inputs the test set into the butterfly segmentation network with loaded weights, segments the butterfly ecological image, and outputs the final butterfly mask image.
[0036] Taking the data set provided by the document "Ecological Photo Dataset for Automatic Identification of Butterfly Species" as an example, the data set has a total of 721 butterfly ecological images, covering 94 species of butterflies, each of which has at least 1 sample and a maximum of 61 samples. The butterfly ecological image segmentation model based on SAM in this embodiment consists of the following steps: (1) Constructing a butterfly segmentation dataset Resize the butterfly image to an image with a height of H pixels and a width of W pixels, marked as , C represents the three channels of the image. The butterfly mask is annotated according to the butterfly boundary using labelimg software to construct a butterfly segmentation dataset, which is divided into a training set and a test set at a ratio of 1:1; (2) Constructing a butterfly segmentation network Figure 2 The structural diagram of the butterfly segmentation network is given. Figure 2In the embodiment, the butterfly segmentation network is composed of an image encoder, a hint encoder, a mask decoder, a two-way convolution module, and a feature fusion module. The image encoder is connected to the two-way convolution module and the mask decoder, and the feature fusion module is connected to the two-way convolution module and the mask decoder. Figure 3 A schematic diagram of the structure of the dual-path convolution module of this embodiment is given. Figure 3 In , the two-way convolutional module consists of a global path and a local path; Figure 4 A schematic diagram of the structure of the global path of this embodiment is given. Figure 4 In , the global pathway is composed of four basic convolutional modules connected sequentially; Figure 5 A schematic diagram of the structure of the local path of this embodiment is given. Figure 5 In , the local path is composed of 8 basic convolutional modules connected sequentially; Figure 6 A schematic diagram of the structure of the basic convolution module of this embodiment is given. Figure 6 In the embodiment, the basic convolution module is composed of a convolution layer with a size of 3*3 and a step size of 1, a normalization layer, a relu activation function layer, a convolution layer with a size of 3*3 and a step size of 1, a normalization layer, and a relu activation function layer. Figure 7 A schematic diagram of the structure of the feature fusion module of this embodiment is given. Figure 7 In the embodiment, the feature fusion module is composed of a 1*256-dimensional vector and a three-layer perceptron.
[0037] Since the present invention adopts a dual-path convolution module and a feature fusion module, the model's feature extraction capability for butterfly images is enhanced, the existing problem of difficulty in extracting butterfly image features is solved, the accuracy of butterfly segmentation is improved, and the parameters of the image encoder, prompt encoder and mask decoder are frozen during the training process, reducing the parameters required for model training, thereby reducing the model training cost.
[0038] (3) Loss Function Loss Function Constructed as follows;
[0039] in Represents the binary cross entropy loss function BCE (Binary CrossEntropy Loss), which is determined by the following formula;
[0040] in, is a specific pixel point, Pixel The probability of predicting the butterfly class, for Corresponding to the true label, is the total number of pixels in an image; represents the Dice loss function, which is determined by the following formula;
[0041] in, is a specific pixel point, Pixel The probability of predicting the butterfly class, for Corresponding to the true label, is the total number of pixels in an image; (4) Training the butterfly segmentation network The training set is input into the butterfly detection network for training. The server graphics card used for training is RTX 3090, the initial learning rate is 0.001, the number of training rounds is 20, the training batch size is 4, and the training is carried out until the loss function converges; (5) Save the model During the training of the butterfly detection network, the weight file after convergence is retained.
[0042] (6) Testing the butterfly segmentation network Input the test set into the trained butterfly segmentation network, load the weight file for testing, segment the butterfly image, output the butterfly mask, and compare it with the real information to verify the performance of the method.
[0043] Complete the butterfly segmentation method based on SAM.
[0044] The method of this embodiment was used to conduct a computer simulation experiment on 6 butterflies. The experimental results are shown in Figure 8 .
[0045] exist Figure 8 In FIG. 1 , (a) to (f) all represent pictures of butterfly ecology in the wild environment in the test set detected by the embodiment method. Figure 8 It can be seen that the method of the present invention has a good effect on the segmentation of butterfly species in the wild environment, with smooth edges, and butterfly pixels can be completely obtained.
[0046] In order to verify the effectiveness of the method of the present invention, a comparative experiment was conducted using the butterfly segmentation method based on SAM in the embodiment and the SAM, SAM-HQ, MobileSAM, FastSAM, EfficientSAM, TinySAM, and MSA segmentation methods to compare the performance of each segmentation method. IoU (Intersection over Union) and MIoU (Mean Intersection over Union) were used as evaluation indicators. The experimental results are shown in Table 1.
[0047] Table 1 Experimental results of the method of this embodiment and the comparative segmentation method (%)
[0048] It can be seen from Table 1 that the IoU and MIoU evaluation indicators of the detection method of the present invention are both optimal values, which are higher than those of the other 7 comparative experimental methods.
[0049] On the other hand, the present invention provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the SAM-based butterfly ecological image segmentation method described in the present invention can be implemented.
[0050] The present invention can also provide a computer device, including a processor and a memory, the memory is used to store a computer executable program, the processor reads the computer executable program from the memory and executes it, and when the processor executes the computer executable program, the SAM-based butterfly ecological image segmentation method described in the present invention can be implemented.
[0051] The computer device may be a laptop computer, a desktop computer or a workstation.
[0052] The processor can be a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or an off-the-shelf field-programmable gate array (FPGA).
[0053] The memory of the present invention may be an internal storage unit of a laptop computer, a desktop computer or a workstation, such as a memory or a hard disk; or an external storage unit, such as a mobile hard disk or a flash memory card.
[0054] Computer-readable storage media may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer-readable storage media may include: read-only memory (ROM), random access memory (RAM), solid-state drive (SSD) or optical disk, etc. Among them, random access memory may include resistance random access memory (ReRAM) and dynamic random access memory (DRAM).
[0055] The present invention is described by way of embodiments, and various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention for those skilled in the art. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the protection scope of the present invention.
[0056] The above contents are only for explaining the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.
Claims
1. A butterfly ecological image segmentation method based on SAM, characterized in that: The following steps are involved: The resized butterfly ecological image is annotated using a mask to obtain a binary butterfly ecological image; The original butterfly ecological image is input into the butterfly segmentation network; the butterfly segmentation network includes an image encoder, a prompt encoder, an image decoder, a two-way convolution module, and a feature fusion module; the image encoder is a pre-trained image encoder of the SAM model, which is composed of 12 Transformer modules connected in sequence, and is used to encode the butterfly ecological image, input the butterfly ecological image, and input the image feature vector output by the last Transformer module into the image decoder; the prompt encoder is a pre-trained prompt encoder of the SAM model, which is used to encode the point, text, and anchor box prompts into a prompt vector; the image decoder is a pre-trained image decoder of the SAM model, which is used to decode the butterfly ecological image feature vector generated by the image encoder and the prompt vector generated by the prompt encoder to generate a mask feature vector; the two-way convolution module is used to further extract features from the feature coding vector generated by the 12 Transformer modules of the image encoder, including two pathways, namely the global pathway and the local pathway. The global pathway uses the image feature vectors generated by the 3rd, 6th, 9th, and 12th Transformer modules of the image encoder to generate an image global feature vector, and the local pathway uses the image feature vectors generated by the remaining Transformer modules to generate an image local feature vector; The global feature vector, local feature vector and mask feature vector of the image are fused to obtain the final butterfly mask image.
2. The butterfly ecological image segmentation method based on SAM according to claim 1 is characterized in that: The resized butterfly ecological image is annotated using a mask, and the binary butterfly ecological image obtained includes: The butterfly ecological image is adjusted to an image with a height of H and a width of W. C represents the image channel. The number of color RGB image channels is 3. The butterfly ecological image is recorded as , use mask method to mark, use labelimg software to make several points on the butterfly boundary, connect the several points with line segments to form a closed image, the closed image takes value 1 inside and value 0 outside, and obtains a png image composed of 0 and 1, where the area with value 1 is the butterfly area.
3. The butterfly ecological image segmentation method based on SAM according to claim 1, characterized in that: When training the butterfly segmentation network: freeze the parameters of the image encoder, prompt encoder, and image decoder, and only train the parameters of the two-way convolution module and the feature fusion module; the loss function used for training the parameters of the two-way convolution module and the feature fusion module is: in Represents the binary cross entropy loss function BCE (Binary CrossEntropy Loss), which is determined by the following formula; in, is a specific pixel point, Pixel The probability of predicting the butterfly class, for Corresponding to the true label, is the total number of pixels in an image; is the Dice loss function, which is determined by the following formula: in, is a specific pixel point, Pixel The probability of predicting the butterfly class, for Corresponding to the true label, is the total number of pixels in an image.
4. The butterfly ecological image segmentation method based on SAM according to claim 1 is characterized in that: The global path is connected in sequence by 4 basic convolution modules, the image feature vector generated by the 3rd Transformer module of the image encoder is input into the 1st basic convolution module of the global path, the image feature vector output by the 1st basic convolution module is added to the image feature vector generated by the 6th Transformer module of the image encoder, and then input into the 2nd basic convolution module of the global path, the image feature vector output by the 2nd basic convolution module is added to the image feature vector generated by the 9th Transformer module of the image encoder, and then input into the 3rd basic convolution module of the global path, the image feature vector output by the 3rd basic convolution module is added to the image feature vector generated by the 12th Transformer module of the image encoder, and then input into the 4th basic convolution module of the global path, and the image feature vector generated by the 4th basic convolution module of the global path is the image global feature vector.
5. The butterfly ecological image segmentation method based on SAM according to claim 1 is characterized in that: The local pathway is connected in sequence by 8 basic convolution modules. The image feature vector generated by the first Transformer module of the image encoder is input into the first basic convolution module of the local pathway. The image feature vector output by the first basic convolution module is added to the image feature vector generated by the second Transformer module of the image encoder, and then input into the second basic convolution module of the local pathway. The image feature vector output by the second basic convolution module is added to the image feature vector generated by the fourth Transformer module of the image encoder, and then input into the third basic convolution module of the local pathway. The image feature vector output by the third basic convolution module is added to the image feature vector generated by the fifth Transformer module of the image encoder, and then input into the fourth basic convolution module of the local pathway. The image feature vector of is added to the image feature vector generated by the 7th Transformer module of the image encoder, and then input into the 5th basic convolution module of the local pathway. The image feature vector output by the 5th basic convolution module is added to the image feature vector generated by the 8th Transformer module of the image encoder, and then input into the 6th basic convolution module of the local pathway. The image feature vector output by the 6th basic convolution module is added to the image feature vector generated by the 10th Transformer module of the image encoder, and then input into the 7th basic convolution module of the local pathway. The image feature vector output by the 7th basic convolution module is added to the image feature vector generated by the 11th Transformer module of the image encoder, and then input into the 8th basic convolution module of the local pathway. The generated image feature vector is the image local feature vector.
6. The butterfly ecological image segmentation method based on SAM according to claim 4 or 5, characterized in that: The basic convolution module is used to extract features from the image feature vector generated by the image encoder Transformer module, and is composed of a convolution layer with a size of 3*3 and a step size of 1, a normalization layer, a relu activation function layer, a convolution layer with a size of 3*3 and a step size of 1, a normalization layer, and a relu activation function layer.
7. The butterfly ecological image segmentation method based on SAM according to claim 1 is characterized in that: The feature fusion module consists of a learnable 1*256-dimensional vector and a three-layer perceptron. The 1*256-dimensional vector is processed by the three-layer perceptron to become a 1*32-dimensional vector. The global feature vector, the local feature vector and the mask vector generated by the two-way convolution module are added together, and the 1*32-dimensional vector is used to perform matrix multiplication with them to obtain the final segmented butterfly mask image.
8. A butterfly ecological image segmentation system based on SAM, characterized in that: It includes image preprocessing module, feature extraction module and feature fusion module; The image preprocessing module is used to mark the resized butterfly ecological image using a mask to obtain a binary butterfly ecological image; The feature extraction module is used to input the original butterfly ecological image into the butterfly segmentation network; the butterfly segmentation network includes an image encoder, a prompt encoder, an image decoder, a two-way convolution module, and a feature fusion module; the image encoder is a pre-trained image encoder of the SAM model, which is composed of 12 Transformer modules connected in sequence, and is used to encode the butterfly ecological image, input the butterfly ecological image, and input the image feature vector output by the last Transformer module into the image decoder; the prompt encoder is a pre-trained prompt encoder of the SAM model, which is used to encode the point, text, and anchor box prompts into a prompt vector; the image decoder is a pre-trained image decoder of the SAM model, which is used to decode the butterfly ecological image feature vector generated by the image encoder and the prompt vector generated by the prompt encoder to generate a mask feature vector; the two-way convolution module is used to further extract features from the feature coding vector generated by the 12 Transformer modules of the image encoder, including two pathways, namely the global pathway and the local pathway. The global pathway uses the image feature vectors generated by the 3rd, 6th, 9th, and 12th Transformer modules of the image encoder to generate an image global feature vector, and the local pathway uses the image feature vectors generated by the remaining Transformer modules to generate an image local feature vector; The feature fusion module is used to fuse the image global feature vector, local feature vector and mask feature vector to obtain the final butterfly mask image.
9. A computer device, characterized in that: It includes a processor and a memory, the memory is used to store a computer executable program, the processor reads part or all of the computer executable program from the memory and executes it, and when the processor executes part or all of the computer executable program, the butterfly ecological image segmentation method based on SAM described in any one of claims 1 to 7 can be implemented.
10. A computer-readable storage medium, characterized in that: A computer program is stored in a computer-readable storage medium, and when the computer program is executed by a processor, it can implement the butterfly ecological image segmentation method based on SAM as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Remote sensing image semantic segmentation method based on multi-scale feature fusion and SAM
CN116206112A
Retinal vessel segmentation method based on gated axial self-attention double-coding convolutional neural network
CN116309629A
Lung CT image segmentation method based on global-local feature correlation fusion
CN116797609A
Remote sensing image building extraction method and device based on double-path jump attention mechanism
CN117475307A
Agent Swin Transform-combined dual-encoder remote sensing image road extraction algorithm
CN118864834A