A semantic segmentation method for farmland drainage ditch remote sensing images based on improved VM-UNet model

Through the improved VM-UNet model, combined with the SENet attention mechanism and multi-scale attention aggregation module, the segmentation problem of optical remote sensing image when the ground features are complex and diverse, achieving a higher precision farmland drainage ditches segmentation effect.

CN119625327BActive Publication Date: 2025-05-06ANHUI AGRICULTURAL UNIVERSITY

Patent Information

Application Number
CN202510168773.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-06
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

When the characteristics of optical remote sensing images are complex and diverse, it is difficult to accurately focus on key areas and important features, and there is a lack of attention to guide cross-level information fusion, resulting in broken segmentation results and confusion in semantic hierarchy.

Method used

Using the improved VM-UNet model, the embedded features are obtained through the Patch Embedding layer, and the SENet attention mechanism and multi-scale attention aggregation module are introduced in the encoder and decoder to perform feature extraction and information fusion, and finally generate semantic segmented images through the projection layer.

Benefits of technology

Achieve higher-precision farmland drainage ditches segmentation in complex farmland scenarios, adapting to more complex farmland scenarios, and improving the model's ability to segment small and narrow targets and distinguish complex ditch backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625327B_ABST
    Figure CN119625327B_ABST
Patent Text Reader

Abstract

The present invention discloses a semantic segmentation method for remote sensing images of farmland drainage ditches based on an improved VM-UNet model. In the method, the VM-UNet model is improved. In the decoder: the features output by the third coding layer and the first coding layer are processed by the SENet attention mechanism and then input to the decoder through a jump connection; the features output by the second coding layer and the Patch Embedding layer are aggregated by multi-scale attention and then input to the decoder through a jump connection; the present invention aggregates spatial features through multi-scale convolution and spatial attention mechanism, enhances the model's ability to resolve complex ditch backgrounds, strengthens the semantic information related to drainage ditches through a channel weighting mechanism, and reduces the influence of confusing background on segmentation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a semantic segmentation method of farmland drainage ditch remote sensing images based on an improved VM-UNet model. Background Art

[0002] In the prior art, for example, CN114549405A discloses a high-resolution remote sensing image semantic segmentation method based on a supervised self-attention network, which uses DialatedResNet as the backbone network, removes the maximum pooling downsampling layer in the last two stages, and replaces it with a hole convolution, and includes a 1×1 convolution layer, a batch normalization layer, and a nonlinear activation layer. After the remote sensing image is input, the features are extracted through the DialatedResNet backbone network, and the hole convolution expands the receptive field without reducing the spatial resolution of the feature map to obtain context information. The 1×1 convolution layer is used to reduce the number of feature channels output by the backbone network and reduce the amount of calculation of subsequent modules. The batch normalization layer accelerates network convergence, and the nonlinear activation layer increases the nonlinear expression ability of the network, and finally outputs the basic feature map. The channel self-attention module includes feature dimension transformation and matrix multiplication operations. In the structure of the category supervised self-attention module, the channel is first reduced through a 1×1 convolution layer, and then through upsampling, softmax layer and other operations, including feature dimension transformation and matrix multiplication. In the spatial supervision self-attention module structure, it is constructed by feature dimension transformation and matrix multiplication, and supervision is performed by constructing an ideal spatial similarity matrix by downsampling the real label. In the advanced feature fusion module structure, 1×1 convolution layer, batch normalization layer and nonlinear activation layer are used to process, and the maximum value is taken to determine the category to which the pixel belongs, so as to achieve the final semantic segmentation. The self-attention mechanism in this patent has high computational complexity and some information is redundant. There may be a lack of supervision when the self-attention mechanism constructs feature context and category information, resulting in redundant and messy information acquisition, reduced semantic segmentation accuracy, confusion between segmented objects and scene context, inaccurate segmentation boundaries, and category misjudgment. It is not suitable for high-precision segmentation tasks.

[0003] For example, CN118365882A discloses an optical remote sensing image segmentation method based on the VMamba model, comprising the following steps: collecting an optical remote sensing image of the scene to be segmented and its corresponding ground object real label map, segmenting the optical remote sensing image into multiple 512×512 data slices, and dividing them into a training set and a validation set according to an 8:2 ratio. Constructing an improved VM-UNet model design: Based on the VMamba model, an asymmetric U-shaped VM-UNet network model is constructed, which innovatively introduces the visual state space (VSS) block as the basic unit. The encoder contains four cascaded VSSLayer layers, the first three layers contain two VSSblock blocks and one PatchMerging2D block, and the fourth layer contains only two VSSblock blocks. The decoder contains four cascaded VSSLayer_up layers, the first layer contains two VSSblock blocks, the last three layers contain two VSSblock blocks and one PatchExpand2D block, and the PatchExpand2D block is used to upsample and restore the image resolution and fuse multi-stage features. Finally, the model training and optimization process training phase parameters are carried out. The patent improves the VM-UNet network model, but it is difficult to accurately focus on key areas and important features due to the complex and diverse features of objects in optical remote sensing images. At the same time, the lack of attention-guided cross-level information fusion may lead to fragmented segmentation results and confusion of semantic levels. Summary of the invention

[0004] In view of the problems existing in the above-mentioned prior art, the present invention provides a semantic segmentation method of farmland drainage ditch remote sensing images based on an improved VM-UNet model. The technical solution is as follows:

[0005] In the first aspect, a semantic segmentation method of farmland drainage ditch remote sensing images based on an improved VM-UNet model is provided, comprising the following steps:

[0006] Input the farmland drainage ditch remote sensing image into the Patch Embedding layer to obtain embedded features;

[0007] Inputting the embedded features into an encoder to obtain an encoding feature map, wherein the encoder includes a first encoding layer, a second encoding layer, a third encoding layer, and a fourth encoding layer connected in sequence;

[0008] Input the encoded feature map into the decoder to obtain a decoded feature map;

[0009] The decoded feature map is input into the projection layer to restore the size of the decoded feature map, and the semantic segmentation image of the farmland drainage ditch remote sensing image is obtained;

[0010] Among them, in the decoder: the third encoding feature output by the third encoding layer is processed by the SENet attention mechanism, and then spliced ​​with the first decoding feature output by the first decoding layer through a jump connection and input into the second decoding layer; the second encoding feature output by the second encoding layer is spliced ​​with the second decoding feature output by the second decoding layer through a jump connection after multi-scale attention aggregation and then input into the third decoding layer; the first encoding feature output by the first encoding layer is processed by the SENet attention mechanism, and then spliced ​​with the third decoding feature output by the third decoding layer through a jump connection and input into the fourth decoding layer; the embedding feature output by the PatchEmbedding layer is spliced ​​with the decoding feature map output by the fourth decoding layer through a jump connection after multi-scale attention aggregation and then input into the projection layer.

[0011] In some embodiments, the multi-scale attention aggregation process includes:

[0012] Through multi-scale fusion and pooling convolution operations, spatial aggregation is performed to obtain the feature map ;

[0013] Through global average pooling operation and convolution, channel aggregation is performed to obtain the feature map ;

[0014] The feature maps processed by spatial aggregation and channel aggregation are added and fused with the original input feature maps respectively to output the final feature map.

[0015] In some embodiments, the spatial aggregation process includes:

[0016] Input feature map , use the input feature map The convolution will have the number of channels Down to the number of channels ,Right now: , ;in, yes Convolution kernel, is the channel compression ratio coefficient, is the processed feature map, For the picture height, is the image width;

[0017] Capture multi-scale spatial information through convolution kernels of different scales: , , ;in, , , Represent convolution kernels of different sizes respectively;

[0018] Add the results of multi-scale convolution element-wise: ;in, It is the element-by-element sum of the results of multi-scale convolution. , , They are the feature maps after convolution with different convolution kernels;

[0019] Add the results of multi-scale convolution element by element Perform mean pooling and maximum pooling, and pass the two through Convolution is performed for aggregation, and the sigmoid activation function is used to generate a spatial attention map; , , ;in, is the feature map after mean pooling, subscript represents the mean, () is the mean pooling operation, is the feature map after the maximum pooling, subscript Indicates the maximum value, () is the maximum pooling operation, is the generated spatial attention map, subscript represents spatial attention, is the Sigmoid activation function;

[0020] Multiply the spatial attention map by the multi-scale features element by element to obtain the spatial attention weighted feature map : .

[0021] In some embodiments, the convolution kernel sizes used in multi-scale convolution are , , ,Right now , , Respectively indicate the size , , The convolution kernel.

[0022] In some embodiments, the process of channel polymerization includes:

[0023] Input feature map , perform global average pooling on the input feature map to obtain the channel dimension features ; ;

[0024] Through two Convolution and ReLU activation functions generate channel attention maps And expand to the same feature map as the input feature ; , ;in, is the channel attention map, subscript Indicates the channel, is the expanded feature map, the superscript Indicates expansion, and is the weight parameter, is the Sigmoid activation function, Indicates the use of ReLU activation function, To expand the operation for attention.

[0025] In some implementations, the feature maps processed by spatial aggregation and channel aggregation are respectively added and fused with the original input feature maps to output the final feature maps, including:

[0026] After the feature map processed by spatial aggregation and channel aggregation is multiplied element by element, it is added and fused with the original input feature map to output the final feature map. )+ ,in, is the feature map after spatial attention weighting, is the feature map of the channel attention machine, and P is the input feature map of the MSAA module.

[0027] In some embodiments, the processing of the SENet attention mechanism includes:

[0028] Input feature map In the compression operation, global average pooling is used to compress the spatial information of the input feature map into a vector of channel dimension. ; ;in, Indicates the number of channels , Tugao , image width , Represents the input feature map In the The first The value of the position, For the Global statistics of channels;

[0029] Nonlinear interaction between channels is achieved through two fully connected layers: , , ;in, is the first fully connected layer, which is used to convert the number of channels from Dimensionality reduction to , is the second fully connected layer, which is used to convert the number of channels from Restore to , is the ReLU activation function, is the Sigmoid activation function;

[0030] Perform channel weighting and generate attention weights Remap to input feature map , to achieve channel-level weighting: , ,in, No. The attention weight of each channel, Input feature map channels, The weighted channel features, and finally output feature maps .

[0031] In some embodiments, the training process of the improved VM-UNet model includes:

[0032] Collect remote sensing image data of farmland drainage ditches and divide the data set into training set, validation set and test set in a ratio of 6:2:2;

[0033] Data enhancement and data annotation were performed. The images of the training set, validation set, and test set were rotated by 0°, 90°, 180°, and 270°. The enhanced images were annotated using the labelme annotation software to obtain a dataset containing two types of labels: drainage ditch and background. The mask image was output in png format. The label of each pixel was the category number, 0 for background and 1 for drainage ditch.

[0034] The data set is input into the improved VM-UNet model to be trained, and the improved VM-UNet model to be trained is trained based on the error loss function between the model output and the actual label to obtain the semantic segmentation model of the farmland drainage ditch remote sensing image;

[0035] After each training cycle, the validation set is used to evaluate the model performance using cross-validation, and the model parameters are adjusted based on the evaluation results.

[0036] In a second aspect, a semantic segmentation model of farmland drainage ditch remote sensing images based on an improved VM-UNet model is provided based on the segmentation method described in the first aspect, including:

[0037] The Patch Embedding layer, encoder, decoder, and projection layer are connected in sequence;

[0038] Among them, the encoder includes a first encoding layer, a second encoding layer, a third encoding layer, and a fourth encoding layer connected in sequence; the decoder includes a first decoding layer, a second decoding layer, a third decoding layer, and a fourth decoding layer connected in sequence; in the decoder: the third encoding feature output by the third encoding layer is processed by the SENet attention mechanism, spliced ​​with the first decoding feature output by the first decoding layer through a jump connection, and then input into the second decoding layer; the second encoding feature output by the second encoding layer is spliced ​​with the second decoding feature output by the second decoding layer through a jump connection after multi-scale attention aggregation, and then input into the third decoding layer; the first encoding feature output by the first encoding layer is processed by the SENet attention mechanism, spliced ​​with the third decoding feature output by the third decoding layer through a jump connection, and then input into the fourth decoding layer; the embedded feature output by the Patch Embedding layer is spliced ​​with the decoding feature map output by the fourth decoding layer through a jump connection after multi-scale attention aggregation, and then input into the projection layer.

[0039] In a third aspect, a computer device is provided, comprising a processor and a memory, wherein the processor executes the segmentation method as described in the first aspect when running computer instructions stored in the memory.

[0040] In a fourth aspect, a computer-readable storage medium is provided, comprising instructions, which, when executed on a computer, enable the computer to execute the segmentation method as described in the first aspect.

[0041] The present invention provides a method for semantic segmentation of remote sensing images of farmland drainage ditches based on an improved VM-UNet model, which has the following beneficial effects: the method for semantic segmentation of remote sensing images of farmland drainage ditches provided by the present invention can achieve higher-precision segmentation of farmland drainage ditches in farmland scenes where the drainage ditches are narrow, long and small in shape, have large scale variations and have complex and changeable backgrounds, and can adapt to more complex farmland scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a structural schematic diagram of a semantic segmentation model for remote sensing images of farmland drainage ditches based on an improved VM-UNet model provided in an embodiment of the present application;

[0043] Figure 2 It is a structural diagram of a multi-scale attention aggregation MSAA module provided in an embodiment of the present application;

[0044] Figure 3 It is a structural diagram of the SENet attention mechanism module provided in an embodiment of the present application;

[0045] Figure 4 It is a flowchart of a training method for a semantic segmentation model of a remote sensing image of a farmland drainage ditch provided in an embodiment of the present application;

[0046] Figure 5 This is a schematic diagram of remote sensing images of farmland drainage ditches;

[0047] Figure 6 The unimproved VM-Unet model is used for Figure 5 Schematic diagram of the segmentation result;

[0048] Figure 7 The improved VM-Unet model is used to Figure 5 Schematic diagram of the segmentation results. DETAILED DESCRIPTION

[0049] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0050] See also Figure 1 , which is a semantic segmentation model of farmland drainage ditch remote sensing images provided by the present invention, is improved based on the VM-UNet model, and the model includes:

[0051] The Patch Embedding layer, encoder, decoder, and projection layer are connected in sequence;

[0052] The encoder includes a first encoding layer, a second encoding layer, a third encoding layer, and a fourth encoding layer connected in sequence; the decoder includes a first decoding layer, a second decoding layer, a third decoding layer, and a fourth decoding layer connected in sequence; in the decoder: the third encoding feature output by the third encoding layer is processed by the SENet attention mechanism module, and then spliced ​​with the first decoding feature output by the first decoding layer through a jump connection and then input to the second decoding layer; the second encoding feature output by the second encoding layer is spliced ​​with the second decoding feature output by the second decoding layer through a jump connection after passing through a multi-scale attention aggregation MSAA (Multi-scale Attention Selection) module, and then input to the third decoding layer; the first encoding feature output by the first encoding layer is processed by the SENet attention mechanism module, and then spliced ​​with the third decoding feature output by the third decoding layer through a jump connection and then input to the fourth decoding layer; the embedded feature output by the Patch Embedding layer is spliced ​​with the decoding feature map output by the fourth decoding layer through a jump connection after passing through the multi-scale attention aggregation module, and then input to the projection layer;

[0053] Among them, in the encoder, the first coding layer, the second coding layer, the third coding layer, and the fourth coding layer respectively use 2 VSS blocks (visual state space modules), and the first coding layer, the second coding layer, and the third coding layer respectively use patch merging operations (Patch Merging layer, used for image block downsampling) at the end; in the decoder, the first decoding layer, the second decoding layer, and the third decoding layer respectively use 2 VSS blocks, the fourth decoding layer uses 1 VSS block, and the second decoding layer, the third decoding layer, and the fourth decoding layer respectively use patch expansion operations (Patch Expanding layer, used for image block upsampling) at the end.

[0054] Based on the above segmentation model, the semantic segmentation method of farmland drainage ditch remote sensing images based on the improved VM-UNet model provided in the embodiment of the present application includes the following steps:

[0055] Step A1: Input the farmland drainage ditch remote sensing image into the Patch Embedding layer to obtain embedded features. The embedded features are input into the first encoding layer of the encoder on the one hand, and input into the end of the decoder through a skip connection on the other hand to be fused with the output of the decoder;

[0056] Step A2, inputting the embedded features into the encoder to obtain a coding feature map, wherein the encoder includes a first coding layer, a second coding layer, a third coding layer, and a fourth coding layer connected in sequence, and respectively obtaining the first coding feature, the second coding feature, the third coding feature, and the fourth coding feature output by the first coding layer, the second coding layer, the third coding layer, and the fourth coding layer. In the encoder, the first coding feature output by the first coding layer is input to the second coding layer, and the first coding feature output by the first coding layer is input to the decoder through a jump connection; the second coding feature output by the second coding layer is input to the third coding layer and the decoder, respectively, and the third coding feature output by the third coding layer is input to the fourth coding layer and the decoder, respectively;

[0057] Step A3, inputting the encoded feature map into a decoder to obtain a decoded feature map;

[0058] Step A4, inputting the decoded feature map into the projection layer to restore the size of the decoded feature map, and obtaining a semantic segmentation image of the farmland drainage ditch remote sensing image;

[0059] Among them, in the decoder: the third encoding feature output by the third encoding layer is processed by the SENet attention mechanism, and then spliced ​​with the first decoding feature output by the first decoding layer through a jump connection and input to the second decoding layer; the second encoding feature output by the second encoding layer is spliced ​​with the second decoding feature output by the second decoding layer through a jump connection and input to the third decoding layer after being processed by the SENet attention mechanism; the first encoding feature output by the first encoding layer is spliced ​​with the third decoding feature output by the third decoding layer through a jump connection and input to the fourth decoding layer after being processed by the SENet attention mechanism; the embedded feature output by the Patch Embedding layer is spliced ​​with the decoding feature map output by the fourth decoding layer through a jump connection and input to the projection layer after being processed by the multi-scale attention aggregation;

[0060] Among them, in the encoder, the first coding layer, the second coding layer, the third coding layer, and the fourth coding layer use 2 VSS blocks respectively, and the first coding layer, the second coding layer, and the third coding layer use patch merging operations at the end respectively; in the decoder, the first decoding layer, the second decoding layer, and the third decoding layer use 2 VSS blocks respectively, the fourth decoding layer uses 1 VSS block, and the second decoding layer, the third decoding layer, and the fourth decoding layer use patch expansion operations at the end respectively.

[0061] In the embodiment of the present application, MSAA aggregates spatial features through multi-scale convolution and spatial attention mechanism, which can effectively capture contextual information of different scales and highlight important areas in space. The jump connection of the first layer of the VM-Unet model mainly contains low-level features such as fine-grained edges and textures, with high resolution and rich detail information. Introducing MSAA in the first layer can capture detail features of different scales, effectively capture edge information of drainage ditches at different scales, and enhance the segmentation ability of small, narrow and long targets; at the same time, it adaptively pays attention to important spatial areas and suppresses background (vegetation, soil, etc.) noise. The feature map resolution of the third layer of the VM-Unet model is further reduced, but it contains detail information and part of the semantic information at the same time. It is a transition layer connecting the shallow layer and the deep layer. The introduction of MSAA can balance the detail information of the shallow layer and the semantic information of the deep layer, capture drainage ditch information of different scales, and make the features have better multi-scale representation capabilities. At the same time, it highlights the spatial features of important areas and enhances the model's ability to resolve complex ditch backgrounds.

[0062] The SENet module dynamically learns the importance of each channel through a channel weighting mechanism and strengthens key channel features. The feature map of the second layer of the VM-Unet model has a higher resolution and a larger number of channels, and contains preliminary semantic information and rich local details. Introducing the SENet module in the second layer can highlight the channel features related to the drainage ditch and suppress redundant background feature channels; strengthen the semantic information related to the drainage ditch and reduce the impact of confusing background on the segmentation results. The feature map of the fourth layer of the VM-Unet model has the lowest resolution and the largest number of channels, and mainly contains high-level semantic information. Introducing the SENet module in the fourth layer can highlight the high-level semantic features of the drainage ditch area, while removing redundancy from the channel, providing the decoder with stronger target recognition capabilities.

[0063] In one embodiment, see Figure 2 The structure of the MSAA module is shown in the figure. The processing of the multi-scale attention aggregation MSAA (Multi-scale Attention Selection) of the MSAA module includes:

[0064] Step A301, perform spatial aggregation through multi-scale fusion and pooling convolution operations to obtain a feature map ;

[0065] Step A302: Perform channel aggregation through global average pooling and convolution to obtain a feature map ;

[0066] Step A303, the feature maps processed by spatial aggregation and channel aggregation are respectively added and fused with the original input feature maps, and the final feature maps are output. Specifically, the feature maps processed by spatial aggregation and channel aggregation are element-by-element multiplied, and then added and fused with the original input feature maps, and the final feature maps are output. )+ ,in, is the feature map after spatial attention weighting, is the feature map of the channel attention machine, and P is the input feature map of the MSAA module.

[0067] In the embodiment of the present application, the MSAA attention mechanism combines channel attention and spatial attention, and can simultaneously focus on the channel correlation and important positions in space of the feature map. The MSAA module introduces spatial multi-scale attention at the first and third jump connections, which can effectively fuse the spatial information of different receptive fields and capture fine-grained and global context features. By adding the MSAA module to the jump connection process of the VM-Unet model, the multi-scale information is combined with the dual attention mechanism to enhance the network's ability to express complex farmland scene features.

[0068] In one implementation, in step A301, the spatial aggregation process includes:

[0069] Step A3011, input feature map , use the input feature map The convolution will have the number of channels Down to the number of channels ,Right now: , ;in, yes Convolution kernel, is the channel compression ratio coefficient, is the processed feature map, For the picture height, is the image width;

[0070] Step A3012, capturing multi-scale spatial information through convolution kernels of different scales, the convolution kernel sizes are respectively , , : , , ;in, , , Represent convolution kernels of different sizes, for example, , , Respectively indicate the size , , The convolution kernel of

[0071] Step A3013, adding the results of the multi-scale convolution element by element: ;in, It is the element-by-element sum of the results of multi-scale convolution. , , They are the feature maps after convolution with different convolution kernels;

[0072] Step A3014, adding the result of the multi-scale convolution element by element Perform mean pooling and maximum pooling, and pass the two through Convolution is performed for aggregation, and the sigmoid activation function is used to generate a spatial attention map; , , ;in, is the feature map after mean pooling, subscript represents the mean, () is the mean pooling operation, is the feature map after the maximum pooling, subscript Indicates the maximum value, () is the maximum pooling operation, is the generated spatial attention map, subscript represents spatial attention, is the Sigmoid activation function;

[0073] Step A3015, multiply the spatial attention map by the multi-scale feature element by element to obtain the spatial attention weighted feature map : .

[0074] In one implementation, in the above step A302, the process of channel aggregation includes:

[0075] Step A3021, input feature map , perform global average pooling on the input feature map to obtain the channel dimension features ; ;

[0076] Step A3022, through two Convolution and ReLU activation functions generate channel attention maps And expand to the same feature map as the input feature ; , ;in, is the channel attention map, subscript Indicates the channel, is the expanded feature map, the superscript Indicates expansion, and is the weight parameter, is the Sigmoid activation function, Indicates the use of ReLU activation function, To expand the operation for attention.

[0077] In one embodiment, in the decoder of step A3 above, SENet is introduced in the second layer and the fourth layer skip connection of the VM-UNet model, and the processing process of the SENet attention mechanism includes:

[0078] Step A311, input feature map In the compression operation, global average pooling is used to compress the spatial information of the input feature map into a vector of channel dimension. ; ;in, Indicates the number of channels , Tugao , image width , Represents the input feature map In the The first The value of the position, For the Global statistics of channels;

[0079] Step A312, realize nonlinear interaction between channels through two fully connected layers: , , ;in, is the first fully connected layer, which is used to convert the number of channels from Dimensionality reduction to , is the second fully connected layer, which is used to convert the number of channels from Restore to , is the ReLU activation function, is the Sigmoid activation function;

[0080] Step A313, perform channel weighting and generate the attention weights Remap to input feature map , to achieve channel-level weighting: , ,in, No. The attention weight of each channel, Input feature map channels, The weighted channel features, and finally output feature maps .

[0081] In the embodiment of the present application, the SENet module adaptively weights the input features in the channel dimension through the channel attention mechanism to improve the network feature expression ability. Introducing SENet in the second and fourth layer jump connections of the VM-UNet model can selectively enhance and suppress the feature map in the channel dimension, thereby improving the performance of the model. By adding the SENet module to the VM-Unet model jump connection process, the channel attention mechanism dynamically enhances important channel features, improves the feature fusion effect of the decoder, and improves the expression ability of VM-UNet for input data.

[0082] In one embodiment, see Figure 4 The training process of the improved VM-UNet model described in the embodiment of the present application includes:

[0083] Step 1: collect remote sensing image data of farmland drainage ditches and divide the data set into training set, verification set and test set in a ratio of 6:2:2;

[0084] Step 2: Data enhancement and data annotation. The images of the training set, validation set, and test set are rotated by 0°, 90°, 180°, and 270°, thereby increasing the data set fourfold. The enhanced images are annotated using the labelme annotation software to obtain a data set containing two labels: drainage ditch and background. The mask image is output in png format. The label of each pixel is the category number, 0 represents background, and 1 represents drainage ditch.

[0085] Step 3, inputting the data set into the improved VM-UNet model to be trained, training the improved VM-UNet model to be trained based on the error loss function between the model output and the actual label, and obtaining a semantic segmentation model of the farmland drainage ditch remote sensing image;

[0086] Step 4: Use cross-validation to evaluate model performance using the validation set after each training cycle and adjust model parameters based on the evaluation results.

[0087] Based on the above-mentioned semantic segmentation method of remote sensing images of farmland drainage ditches, the embodiment of the present application further provides a semantic segmentation device of remote sensing images of farmland drainage ditches based on an improved VM-UNet model, comprising:

[0088] The image embedding unit is used to input the remote sensing image of farmland drainage ditches into the Patch Embedding layer to obtain embedded features;

[0089] A coding data acquisition unit, used for inputting the embedded features into an encoder to obtain a coding feature map, wherein the encoder includes a first coding layer, a second coding layer, a third coding layer, and a fourth coding layer connected in sequence;

[0090] A decoding data acquisition unit, used for inputting the encoding feature map into a decoder to obtain a decoding feature map;

[0091] A semantic segmentation result acquisition unit is used to restore the size of the decoded feature map through the projection layer to obtain a semantic segmentation image;

[0092] Among them, in the decoder: the third encoding feature output by the third encoding layer is processed by the SENet attention mechanism, and then spliced ​​with the first decoding feature output by the first decoding layer through a jump connection and input into the second decoding layer; the second encoding feature output by the second encoding layer is spliced ​​with the second decoding feature output by the second decoding layer through a jump connection after multi-scale attention aggregation and then input into the third decoding layer; the first encoding feature output by the first encoding layer is processed by the SENet attention mechanism, and then spliced ​​with the third decoding feature output by the third decoding layer through a jump connection and input into the fourth decoding layer; the embedding feature output by the PatchEmbedding layer is spliced ​​with the decoding feature map output by the fourth decoding layer through a jump connection after multi-scale attention aggregation and then input into the projection layer.

[0093] For the specific limitations of the semantic segmentation device for remote sensing images of farmland drainage ditches based on the improved VM-UNet model, please refer to the limitations of the semantic segmentation method for remote sensing images of farmland drainage ditches based on the improved VM-UNet model in the above text, which will not be repeated here.

[0094] An embodiment of the present application provides a computer device, including a processor and a memory. When the processor runs the computer instructions stored in the memory, it executes the above-mentioned semantic segmentation method of farmland drainage ditch remote sensing images based on the improved VM-UNet model.

[0095] The processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.

[0096] The memory may include one or more computer-readable storage media, which may be non-transitory. The memory may also include high-speed random access memory, and non-volatile memory, such as one or more disk storage devices, flash memory storage devices.

[0097] In some embodiments, the computer device may also optionally include: an input interface and an output interface. The processor, the memory and the input interface and the output interface may be connected via a bus or a signal line. Each peripheral device may be connected to the input interface and the output interface via a bus, a signal line or a circuit board. The input interface and the output interface may be used to connect at least one peripheral device related to input / output to the processor and the memory.

[0098] An embodiment of the present application provides a computer-readable storage medium, including instructions, which, when executed on a computer, enable the computer to execute the steps of the semantic segmentation method of farmland drainage ditch remote sensing images based on the improved VM-UNet model.

[0099] For example, the computer readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage node, etc.

[0100] The present invention is not limited to the above-mentioned specific implementation modes. Various changes made by ordinary technicians in this field based on the above-mentioned concepts without creative work are all within the protection scope of the present invention.

Claims

1. A semantic segmentation method for farmland drainage ditch remote sensing images based on an improved VM-UNet model, characterized in that: The steps include: Input the farmland drainage ditch remote sensing image into the Patch Embedding layer to obtain embedded features; Inputting the embedded features into an encoder to obtain an encoding feature map, wherein the encoder includes a first encoding layer, a second encoding layer, a third encoding layer, and a fourth encoding layer connected in sequence; Input the encoded feature map into the decoder to obtain a decoded feature map; The decoded feature map is input into the projection layer to restore the size of the decoded feature map, and the semantic segmentation image of the farmland drainage ditch remote sensing image is obtained; Among them, in the decoder: the third encoding feature output by the third encoding layer is processed by the SENet attention mechanism, and then spliced ​​with the first decoding feature output by the first decoding layer through a jump connection and input into the second decoding layer; the second encoding feature output by the second encoding layer is spliced ​​with the second decoding feature output by the second decoding layer through a jump connection after multi-scale attention aggregation and then input into the third decoding layer; the first encoding feature output by the first encoding layer is processed by the SENet attention mechanism, and then spliced ​​with the third decoding feature output by the third decoding layer through a jump connection and input into the fourth decoding layer; the embedding feature output by the PatchEmbedding layer is spliced ​​with the decoding feature map output by the fourth decoding layer through a jump connection after multi-scale attention aggregation and then input into the projection layer.

2. The method for semantic segmentation of farmland drainage ditch remote sensing images based on the improved VM-UNet model according to claim 1, characterized in that: The processing process of multi-scale attention aggregation includes: Through multi-scale fusion and pooling convolution operations, spatial aggregation is performed to obtain the feature map ; Through global average pooling operation and convolution, channel aggregation is performed to obtain the feature map ; The feature maps processed by spatial aggregation and channel aggregation are added and fused with the original input feature maps respectively to output the final feature map.

3. The semantic segmentation method of farmland drainage ditch remote sensing image based on the improved VM-UNet model according to claim 2 is characterized in that: The process of spatial aggregation includes: Input feature map , use the input feature map The convolution will have the number of channels Down to the number of channels ,Right now: , ;in, yes Convolution kernel, is the channel compression ratio coefficient, is the processed feature map, For the picture height, is the image width; Capture multi-scale spatial information through convolution kernels of different scales: , , ;in, , , Represent convolution kernels of different sizes respectively; Add the results of multi-scale convolution element-wise: ;in, It is the element-by-element sum of the results of multi-scale convolution. , , They are the feature maps after convolution with different convolution kernels; Add the results of multi-scale convolution element by element Perform mean pooling and maximum pooling, and pass the two through Convolution is performed for aggregation, and the sigmoid activation function is used to generate a spatial attention map; , , ;in, is the feature map after mean pooling, subscript represents the mean, () is the mean pooling operation, is the feature map after the maximum pooling, subscript Indicates the maximum value, () is the maximum pooling operation, is the generated spatial attention map, subscript represents spatial attention, is the Sigmoid activation function; Multiply the spatial attention map by the multi-scale features element by element to obtain the spatial attention weighted feature map : .

4. The method for semantic segmentation of farmland drainage ditch remote sensing images based on the improved VM-UNet model according to claim 2, characterized in that: The process of channel aggregation includes: Input feature map , perform global average pooling on the input feature map to obtain the channel dimension features ; ; Through two Convolution and ReLU activation functions generate channel attention maps And expand to the same feature map as the input feature ; , ;in, is the channel attention map, subscript Indicates the channel, is the expanded feature map, the superscript Indicates expansion, and is the weight parameter, is the Sigmoid activation function, Indicates the use of ReLU activation function, To expand the operation for attention.

5. The method for semantic segmentation of farmland drainage ditch remote sensing images based on the improved VM-UNet model according to claim 2, characterized in that: The feature maps processed by spatial aggregation and channel aggregation are respectively added and fused with the original input feature maps to output the final feature maps, including: After the feature map processed by spatial aggregation and channel aggregation is multiplied element by element, it is added and fused with the original input feature map to output the final feature map. )+ ,in, is the feature map after spatial attention weighting, is the feature map of the channel attention machine, and P is the input feature map of the MSAA module.

6. The method for semantic segmentation of farmland drainage ditch remote sensing images based on the improved VM-UNet model according to claim 1, characterized in that: The processing process of the SENet attention mechanism includes: Input feature map In the compression operation, global average pooling is used to compress the spatial information of the input feature map into a vector of channel dimension. ; ;in, Indicates the number of channels , Tugao , image width , Represents the input feature map In the The first The value of the position, For the Global statistics of channels; Nonlinear interaction between channels is achieved through two fully connected layers: , , ;in, is the first fully connected layer, which is used to convert the number of channels from Dimensionality reduction to , is the second fully connected layer, which is used to convert the number of channels from Restore to , is the ReLU activation function, is the Sigmoid activation function; Perform channel weighting and generate attention weights Remap to input feature map , to achieve channel-level weighting: , ,in, No. The attention weight of each channel, Input feature map channels, The weighted channel features, and finally output feature maps .

7. The method for semantic segmentation of farmland drainage ditch remote sensing images based on the improved VM-UNet model according to claim 1, characterized in that: The training process of the improved VM-UNet model includes: Collect remote sensing image data of farmland drainage ditches and divide the data set into training set, validation set and test set in a ratio of 6:2:2; Data enhancement and data annotation were performed. The images of the training set, validation set, and test set were rotated by 0°, 90°, 180°, and 270°. The enhanced images were annotated using the labelme annotation software to obtain a dataset containing two types of labels: drainage ditch and background. The mask image was output in png format. The label of each pixel was the category number, 0 for background and 1 for drainage ditch. The data set is input into the improved VM-UNet model to be trained, and the improved VM-UNet model to be trained is trained based on the error loss function between the model output and the actual label to obtain the semantic segmentation model of the farmland drainage ditch remote sensing image; After each training cycle, the validation set is used to evaluate the model performance using cross-validation, and the model parameters are adjusted based on the evaluation results.

8. A semantic segmentation model for farmland drainage ditch remote sensing images based on the segmentation method of claim 1 and based on an improved VM-UNet model, characterized in that: include: The Patch Embedding layer, encoder, decoder, and projection layer are connected in sequence; Among them, the encoder includes a first encoding layer, a second encoding layer, a third encoding layer, and a fourth encoding layer connected in sequence; the decoder includes a first decoding layer, a second decoding layer, a third decoding layer, and a fourth decoding layer connected in sequence; in the decoder: the third encoding feature output by the third encoding layer is processed by the SENet attention mechanism, spliced ​​with the first decoding feature output by the first decoding layer through a jump connection, and then input into the second decoding layer; the second encoding feature output by the second encoding layer is spliced ​​with the second decoding feature output by the second decoding layer through a jump connection after multi-scale attention aggregation, and then input into the third decoding layer; the first encoding feature output by the first encoding layer is processed by the SENet attention mechanism, spliced ​​with the third decoding feature output by the third decoding layer through a jump connection, and then input into the fourth decoding layer; the embedded feature output by the Patch Embedding layer is spliced ​​with the decoding feature map output by the fourth decoding layer through a jump connection after multi-scale attention aggregation, and then input into the projection layer.

9. A computer device, characterized in that: The method comprises a processor and a memory, wherein the processor executes the method according to any one of claims 1 to 7 when running the computer instructions stored in the memory.

10. A computer-readable storage medium, characterized in that: The method comprises instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • High-resolution remote sensing image semantic segmentation method based on supervised self-attention network

    CN114549405A

  • Image segmentation method based on boundary enhancement

    CN116205927A

  • Remote sensing image semantic segmentation method based on multi-scale attention and double encoders

    CN117496151A

Cited By

  • Hyperspectral remote sensing image interpretation method based on TeRN improved UNetFormer

    CN121236561A