Method and device for segmenting farmland image by fusing rgb and multispectral, equipment and medium
By extracting and enhancing features from RGB and multispectral images of farmland, and adaptively fusing deep spatial-spectral feature maps, the problem of low accuracy in farmland image segmentation is solved, and high-precision farmland image segmentation is achieved.
Patent Information
- Application Number
- CN202310412428.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-10
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2043-04-10
AI Technical Summary
Existing technologies cannot provide high-precision farmland image segmentation methods, resulting in low accuracy in farmland image segmentation, which makes it difficult to meet the needs of precision agricultural management.
By extracting and enhancing features from RGB and multispectral images of farmland, a deep spatial-spectral feature map is obtained, and adaptive fusion is performed to improve the segmentation accuracy of farmland images.
By dynamically fusing data features from different modalities, semantic information is enriched, and inconsistencies between features from multiple different modalities are suppressed, significantly improving the segmentation accuracy and robustness of farmland segmentation.
Smart Images

Figure CN116524184B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of computer, and particularly relates to a farmland image segmentation method and device fusing RGB and multispectral, equipment and medium. BACKGROUND
[0002] Farmland segmentation is the basis for analyzing farmland environment and farmland use in intelligent agricultural management. Existing research mainly performs semantic segmentation on farmland based on remote sensing hyperspectral images, but the segmentation accuracy is low, which is difficult to meet the requirements of fine agricultural management. Based on the RGB images taken by unmanned aerial vehicles near the ground, the accuracy of farmland semantic segmentation can be improved, but due to the diversity of farmland ecology and the irregularity of farmland boundaries, as well as the influence of light and weather conditions during aerial photography, the accuracy of farmland semantic segmentation based on near-ground RGB images alone also faces great challenges. SUMMARY
[0003] The present application aims to provide a farmland image segmentation method and device fusing RGB and multispectral, equipment and medium, which aims to solve the problem of low farmland image segmentation accuracy due to the fact that the prior art cannot provide an effective farmland image segmentation method.
[0004] In one aspect, the present application provides a farmland image segmentation method fusing RGB and multispectral, which comprises the following steps:
[0005] Feature extraction is performed on the input farmland RGB image to obtain N first feature images, and feature enhancement is performed on the N first feature images to obtain corresponding N first deep spectral feature maps, wherein N is a positive integer greater than 1;
[0006] Feature extraction is performed on the input farmland multispectral image to obtain N second feature images, and feature enhancement is performed on the N second feature images to obtain corresponding N second deep spectral feature maps, wherein the farmland multispectral image is a multispectral image of the farmland RGB image;
[0007] The N first deep spectral feature maps and the N second deep spectral feature maps are adaptively fused to obtain a prediction map of the farmland in the farmland RGB image.
[0008] In another aspect, the present application provides a farmland image segmentation device fusing RGB and multispectral, which further comprises:
[0009] The first feature acquisition unit is configured to perform feature extraction on the input farmland RGB image to obtain N first feature images, and perform feature enhancement on the N first feature images to obtain corresponding N first deep spectral feature maps, wherein N is a positive integer greater than 1;
[0010] The second feature acquisition unit is configured to perform feature extraction on the input farmland multi-spectral image to obtain N second feature images, and perform feature enhancement on the N second feature images to obtain corresponding N second deep spectral feature maps, wherein the farmland multi-spectral image is a multi-spectral image of the farmland RGB image.
[0011] The feature fusion unit is configured to perform adaptive fusion on the N first deep spectral feature maps and the N second deep spectral feature maps to obtain a prediction map of the farmland in the farmland RGB image.
[0012] In another aspect, the present application also provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the above method when executing the computer program.
[0013] In another aspect, the present application also provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the above method when executed by a processor.
[0014] The present application performs feature extraction on the input farmland RGB image to obtain N first feature images, performs feature enhancement on the N first feature images to obtain corresponding N first deep spectral feature maps, performs feature extraction on the input farmland multi-spectral image to obtain N second feature images, performs feature enhancement on the N second feature images to obtain corresponding N second deep spectral feature maps, and performs adaptive fusion on the N first deep spectral feature maps and the N second deep spectral feature maps to obtain a prediction map of the farmland in the farmland RGB image, thereby improving the segmentation accuracy of the farmland image in the farmland RGB image. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 is an implementation flowchart of the farmland image segmentation method provided by the first embodiment of the present application;
[0016] Figure 2a is a structural schematic diagram of the RGB image encoder provided by the second embodiment of the present application;
[0017] Figure 2b is a structural schematic diagram of the first convolutional pooling module provided by the second embodiment of the present application;
[0018] Figure 2c is a structural schematic diagram of the convolutional pooling module provided by the second embodiment of the present application;
[0019] Figure 2d is a structural schematic diagram of the first parallel convolutional module provided by the second embodiment of the present application;
[0020] Figure 2e is a structural schematic diagram of the spectral feature attention module provided in Embodiment Two of the present application;
[0021] Figure 3 is a structural schematic diagram of the multi-spectral image encoder provided in Embodiment Three of the present application;
[0022] Figure 4a is a structural schematic diagram of the decoder provided in Embodiment Four of the present application;
[0023] Figure 4b is a structural schematic diagram of the multi-modal feature adaptive module provided in Embodiment Four of the present application;
[0024] Figure 5 is an implementation flowchart of the farmland image segmentation method provided in Embodiment Five of the present application;
[0025] Figure 6 is a structural schematic diagram of the farmland image segmentation device provided in Embodiment Six of the present application; and
[0026] Figure 7 is a structural schematic diagram of the computer device provided in Embodiment Seven of the present application. DETAILED DESCRIPTION
[0027] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application is further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0028] The specific implementation of the present application is described in detail below in combination with specific embodiments:
[0029] Example One:
[0030] Figure 1 An implementation flowchart of the farmland image segmentation method provided in Embodiment One of the present application is shown, and only the part related to the embodiments of the present application is shown for the convenience of description, and the details are as follows:
[0031] In step S101, the input farmland RGB image is subjected to feature extraction to obtain N first feature images, and the N first feature images are subjected to feature enhancement to obtain corresponding N first deep spectral feature maps;
[0032] The embodiment of the present application is suitable for a computer device for image processing, such as a personal computer or a server, to segment or predict a farmland image from a farmland RGB image, that is, to mark or divide a farmland contour in the image. After receiving a to-be-segmented farmland RGB image including a farmland, feature extraction is performed on the to-be-segmented farmland RGB image to obtain N first feature images, which are deep spectral features with representativeness and discriminativeness. Then, feature enhancement is performed on the N first feature images to selectively enhance useful spectral features and suppress useless features, and finally N corresponding first deep spectral feature maps are obtained, wherein N is a positive integer greater than 1.
[0033] In step S102, feature extraction is performed on the input farmland multispectral image to obtain N second feature images, and feature enhancement is performed on the N second feature images to obtain N corresponding second deep spectral feature maps.
[0034] In the embodiment of the present application, the farmland multispectral image is a multispectral image of the farmland RGB image, that is, a multispectral image obtained by using a spectral camera to capture the same scene as the farmland RGB image. The farmland RGB image and the farmland multispectral image can be captured by one multispectral camera for the same farmland scene, or can be captured by one conventional camera and multiple monochromatic cameras for the same farmland scene. After obtaining the farmland multispectral image corresponding to the farmland RGB image, feature extraction is performed on the farmland multispectral image to obtain N second feature images, and feature enhancement is performed on the N second feature images to obtain N corresponding second deep spectral feature maps.
[0035] In step S103, adaptive fusion is performed on the N first deep spectral feature maps and the second deep spectral feature maps to obtain a prediction map of the farmland in the farmland RGB image.
[0036] In the embodiment of the present application, feature extraction is performed on the input farmland RGB image to obtain N first feature images, feature enhancement is performed on the N first feature images to obtain N corresponding first deep spectral feature maps, N is a positive integer greater than 1, feature extraction is performed on the input farmland multispectral image to obtain N second feature images, feature enhancement is performed on the N second feature images to obtain N corresponding second deep spectral feature maps, the farmland multispectral image is a multispectral image obtained by using a spectral camera to capture the same scene as the farmland RGB image, and adaptive fusion is performed on the N first deep spectral feature maps and the second deep spectral feature maps to obtain a prediction map of the farmland in the farmland RGB image. Thus, the farmland RGB image and the farmland multispectral image are fused together, and the segmentation robustness and the segmentation accuracy of the farmland image in the farmland RGB image are improved.
[0037] Example Two:
[0038] In the embodiment of the application, step S101 in the above embodiment one is implemented by the RGB image encoder. Specifically, the input farmland RGB image is subjected to feature extraction by the N convolutional pooling modules and one parallel convolution module of the RGB image encoder to obtain N first feature images, and the N first feature images are subjected to feature enhancement by the N spectral feature attention modules of the RGB image encoder to obtain corresponding N first deep spectral feature maps.
[0039] In a preferred embodiment, as shown in Figure 2a the value of N is 4, and the RGB image encoder includes a first convolutional pooling module 211, a second convolutional pooling module 212, a third convolutional pooling module 213, a first parallel convolution module 214, a first spectral feature attention module 221, a second spectral feature attention module 222, a third spectral feature attention module 223, and a fourth spectral feature attention module 224. The first convolutional pooling module 211, the second convolutional pooling module 212, and the third convolutional pooling module 213 are used to extract features from the input image to obtain deep spectral features, the first parallel convolution module 214 is used to extract context information from the input image to obtain deep spectral features with rich context information, and the first spectral feature attention module 221, the second spectral feature attention module 222, the third spectral feature attention module 223, and the fourth spectral feature attention module 224 are used to enhance the received feature images to selectively enhance useful spectral features and suppress useless features, and finally obtain corresponding multiple first deep spectral feature maps, thereby improving the feature expression of the extracted farmland RGB image. Specifically, as shown in the figure, the input of the first convolutional pooling module 211 is the farmland RGB image, the output of the first convolutional pooling module 211 is the input of the second convolutional pooling module 212 and the first spectral feature attention module 221, the output of the second convolutional pooling module 212 is the input of the third convolutional pooling module 213 and the second spectral feature attention module 222, the output of the third convolutional pooling module 213 is the input of the first parallel convolution module 214 and the third spectral feature attention module 223, the output of the first parallel convolution module 214 is the input of the fourth spectral feature attention module 224, and the outputs of the first spectral feature attention module 221, the second spectral feature attention module 222, the third spectral feature attention module 223, and the fourth spectral feature attention module 224 are the first deep spectral feature maps.
[0040] In an embodiment, as shown in Figure 2bAs shown, the first convolutional pooling module includes a standard convolutional layer and a pooling layer to perform preliminary downsampling processing on the input RGB farmland image to obtain the corresponding deep spatial spectral features. The standard convolutional layer consists of a 3×3 kernel convolution, batch normalization (BatchNorm), and a LeakReLU activation function. The pooling layer is used to perform pooling operations on the feature map after convolution by the standard convolutional layer. In one embodiment, as... Figure 2c As shown, the second and third convolutional pooling modules each consist of two standard convolutional layers (the first and second standard convolutional layers in the figure) and a pooling layer. These layers further downsample the deep spatial-spectral features output by the first convolutional pooling module to compress the semantic information in the farmland image, thereby obtaining the corresponding deep spatial-spectral features. In one embodiment, as... Figure 2d As shown, the first parallel convolutional module consists of a regular convolutional layer with a kernel size of 3×3, a dilated convolutional layer with a kernel size of 3×3 (also known as dilated convolution), a splicing layer, and a standard convolutional layer, thereby extracting the deep features output by the third convolutional pooling module in parallel, enriching the context space and channel information of the farmland, enhancing feature expression, and obtaining enhanced deep spatial-spectral features.
[0041] In one embodiment, the first, second, third, and fourth spatial spectral feature attention modules have the same structure, such as... Figure 2e As shown, each spatial spectral feature attention module includes first, second, third, and fourth pooling layers, first, second, and third convolutional layers, and first and second activation functions. Specifically, the first pooling layer performs global average pooling on the input feature map, the second pooling layer performs global max pooling on the input feature map, the third pooling layer performs channel-direction-based average pooling on the input feature map, and the fourth pooling layer performs max pooling on the input feature map. The first and second activation functions can be sigmoid activation functions. To capture the correlation between channels after pooling, the first and second convolutional layers perform compressed feature processing on the channel descriptions output by global average pooling and global max pooling, generate channel weights, and then sum them. After feature normalization through the first activation function, the weighted feature map is multiplied by the input feature map to obtain the deep spatial spectral weighted feature map F. c To obtain prominent spatial information from deep spatial spectral features that complement channel importance, the average pooling and max pooling outputs of the third and fourth pooling layers are concatenated based on channel direction to obtain a cascaded feature map of H×W×2. This map is then passed through a third convolutional layer to generate a feature map with spatial weights, which enhances useful spatial locations or suppresses useless ones. Finally, the features are normalized using a second activation function to obtain F. s Finally, to enhance the extracted deep spatial spectral features, F c With F sThe multiplication weighting is performed to obtain the enhanced deep spectral feature. The spectral feature attention module can improve the expression of the farmland feature extracted by the RGB image encoder.
[0042] Example Three:
[0043] In the embodiment, the step S102 in the embodiment one is implemented by the multispectral image encoder. Specifically, the multispectral image encoder is used to extract the features of the input farmland multispectral image through the N convolution pooling modules and the parallel convolution module, to obtain N second feature images, and to enhance the features of the N second feature images through the N spectral feature attention modules of the multispectral image encoder, to obtain corresponding N second deep spectral feature maps.
[0044] In a preferred embodiment, the value of N is 4, as shown in the following figure. Figure 3 The multispectral image encoder includes the fourth convolution pooling module 311, the fifth convolution pooling module 312, the sixth convolution pooling module 313, the second parallel convolution module 314, the fifth spectral feature attention module 321, the sixth spectral feature attention module 322, the seventh spectral feature attention module 323, and the eighth spectral feature attention module 324. Specifically, as shown in the figure, the input of the fourth convolution pooling module is the farmland multispectral image, the output of the fourth convolution pooling module is the input of the fifth convolution pooling module and the fifth spectral feature attention module, the output of the fifth convolution pooling module is the input of the sixth convolution pooling module and the sixth spectral feature attention module, the output of the sixth convolution pooling module is the input of the second parallel convolution module and the seventh spectral feature attention module, and the output of the second parallel convolution module is the input of the eighth spectral feature attention module.
[0045] In specific implementations, the functions and structures of the fourth convolutional pooling module 311, the fifth convolutional pooling module 312, the sixth convolutional pooling module 313, and the second parallel convolutional module 314 of the multispectral image encoder can be the same as or identical to those of the first convolutional pooling module 211, the second convolutional pooling module 212, the third convolutional pooling module 213, and the first parallel convolutional module 214 of the RGB image encoder. For details, please refer to the description in Embodiment 2, which will not be repeated here. The functions and structures of the fifth spatial-spectral feature attention module 321, the sixth spatial-spectral feature attention module 322, the seventh spatial-spectral feature attention module 323, and the eighth spatial-spectral feature attention module 324 of the multispectral image encoder can be the same as or identical to those of the first spatial-spectral feature attention module 221, the second spatial-spectral feature attention module 222, the third spatial-spectral feature attention module 223, and the fourth spatial-spectral feature attention module 224 of the RGB image encoder. For details, please refer to the description in Embodiment 2, which will not be repeated here. Through this spatial-spectral feature attention module of the multispectral image encoder, the feature representation of the extracted farmland multispectral image can be improved.
[0046] Example Four:
[0047] In the embodiment of the invention, step S103 in the above embodiment one is implemented by the decoder. Specifically, the decoder adaptively fuses N first deep spatial spectral feature maps and second deep spatial spectral feature maps to obtain a predicted map of farmland in the farmland RGB image.
[0048] In a preferred embodiment, such as Figure 4a As shown, the value of N is 4, and the decoder includes a first multimodal feature adaptation module 411, a second multimodal feature adaptation module 412, a third multimodal feature adaptation module 413, a fourth multimodal feature adaptation module 414, an upsampling layer 421, a first decoding module 431, a second decoding module 432, and a third decoding module 433. The inputs of the first multimodal feature adaptive module are the first deep spatial-spectral feature map and the second deep spatial-spectral feature map. The output of the first multimodal feature adaptive module is the input of the upsampling layer. The inputs of the second multimodal feature adaptive module are the output of the upsampling layer, the first deep spatial-spectral feature map, and the second deep spatial-spectral feature map. The input of the first decoding module is the output of the second multimodal feature adaptive module. The inputs of the third multimodal feature adaptive module are the output of the first decoding module, the first deep spatial-spectral feature map, and the second deep spatial-spectral feature map. The input of the second decoding module is the output of the third multimodal feature adaptive module. The inputs of the fourth multimodal feature adaptive module are the output of the second decoding module, the first deep spatial-spectral feature map, and the second deep spatial-spectral feature map. The input of the third decoding module is the output of the fourth multimodal feature adaptive module. The output of the fourth decoding module is the prediction map.
[0049] In a specific implementation, such as Figure 4b As shown, the multimodal feature adaptive module includes first, second, third, fourth, fifth, and sixth convolutional layers and a normalized exponential function. When adaptively fusing N first deep spatial spectral feature maps and second deep spatial spectral feature maps, the multimodal feature adaptive module includes the following steps:
[0050] (1) Use the first and second convolutional layers to perform convolution operations on the input first and second deep spatial spectral feature maps respectively to obtain the corresponding first and second intermediate feature maps. The first and second intermediate feature maps have the same number of channels.
[0051] (2) Use the third and fourth convolutional layers to perform convolution operations on the first and second intermediate feature maps respectively, so as to compress the number of channels of the first and second intermediate feature maps and obtain the third and fourth intermediate feature maps;
[0052] (3) Perform channel splicing operation on the third and fourth intermediate feature maps to obtain the fifth intermediate feature map;
[0053] (4) Use the fifth convolutional layer to perform a convolution operation on the fifth intermediate feature map to compress the number of channels of the fifth intermediate feature map and obtain the sixth intermediate feature map;
[0054] (5) Obtain the first and second weight parameters corresponding to the sixth intermediate feature map through the normalized exponential function;
[0055] (6) Based on the first and second weight parameters and the first and second intermediate feature maps, obtain the first and second weighted feature maps, and add the first and second weighted feature maps to obtain the seventh intermediate feature map;
[0056] (7) Use the sixth convolutional layer to perform a convolution operation on the seventh intermediate feature map to restore the number of channels of the seventh intermediate feature map and obtain the fused feature map.
[0057] Through the above steps, the multimodal feature adaptive module can fuse the first deep spatial spectral feature map of the farmland RGB image and the second deep spatial spectral feature map of the farmland multispectral image, thereby dynamically fusing data features between different modalities, enriching semantic information, suppressing inconsistencies between multiple different modal features, and significantly improving the segmentation accuracy of farmland segmentation.
[0058] In an embodiment, the upsampling layer is composed of a bilinear interpolation operation, which is used to interpolate the feature obtained by the first multi-modal feature adaptive module fusion to obtain a larger size of the fused feature to restore the size of the feature map. In an embodiment, the first decoding module is composed of one convolution and one upsampling operation, which is used to add the output of the second multi-modal feature adaptive module and the upsampling layer, and then perform decoding processing to gradually restore the size of the feature map and its information. The second decoding module is composed of one convolution and one upsampling operation, which is used to add the output of the third multi-modal feature adaptive module and the first decoding module, and then perform decoding processing to gradually restore the size of the feature map and its information. The second decoding module is composed of two convolutions and one upsampling operation, which is used to add the output of the fourth multi-modal feature adaptive module and the second decoding module, and then perform decoding processing to output the prediction map.
[0059] The decoder of the embodiment of the present application adaptively fuses the N first deep spectral feature maps and the second deep spectral feature maps to obtain the prediction map of the farmland in the farmland RGB image, thereby dynamically fusing the data features between different modalities, enriching the semantic information, and suppressing the inconsistency between different modalities, and greatly improving the segmentation accuracy of farmland segmentation.
[0060] Example Five:
[0061] Figure 5 The implementation process of the farmland image segmentation method provided by the embodiment five of the present application is shown, only the part related to the embodiment of the present application is shown for convenience of description, and the details are described as follows:
[0062] In step S501, the N convolutional pooling modules and one parallel convolutional module of the RGB image encoder are used to extract features from the input farmland RGB image to obtain N first feature images, and the N spectral feature attention modules of the RGB image encoder are used to enhance the features of the N first feature images to obtain corresponding N first deep spectral feature maps.
[0063] The embodiment of the present application is applicable to a computer device for image processing, such as a personal computer or a server, to segment or predict a farmland image from a farmland RGB image, i.e., to mark or divide the farmland contour in the image. After receiving the to-be-segmented farmland RGB image including the farmland, the RGB image encoder is used to extract features from the to-be-segmented farmland RGB image to obtain N first feature images, and the first feature images are deep spectral features with representativeness and discriminativeness. Then, the N first feature images are enhanced to selectively enhance useful spectral features and suppress useless features, and finally N corresponding first deep spectral feature maps are obtained, wherein N is a positive integer greater than 1, and the specific structure of the RGB image encoder can be referred to the description of the embodiment two, which will not be described here.
[0064] In step S502, the input farmland multispectral image is subjected to feature extraction by N convolution pooling modules and one parallel convolution module of the multispectral image encoder to obtain N second feature images, and the N second feature images are subjected to feature enhancement by N spectral feature attention modules of the multispectral image encoder to obtain corresponding N second deep spectral feature maps;
[0065] In the embodiment of the present application, the farmland multispectral image is a multispectral image obtained by using a multispectral camera to capture the same scene of the farmland RGB image, and the farmland RGB image and the farmland multispectral image can be captured by one multispectral camera for the same farmland scene, or can be captured by one conventional camera and multiple monochromatic cameras for the same farmland scene. After obtaining the farmland multispectral image corresponding to the farmland RGB image, the farmland multispectral image is subjected to feature extraction by the multispectral image encoder to obtain N second feature images, and the N second feature images are subjected to feature enhancement to obtain corresponding N second deep spectral feature maps. The specific structure of the multispectral image encoder can refer to the description of Embodiment Three, and will not be repeated here.
[0066] In step S503, the N first deep spectral feature maps and the second deep spectral feature maps are subjected to adaptive fusion by the decoder to obtain the prediction map of the farmland in the farmland RGB image.
[0067] In the embodiment of the present application, the N first deep spectral feature maps and the second deep spectral feature maps are subjected to adaptive fusion by the decoder to obtain the prediction map of the farmland in the farmland RGB image. The specific structure of the decoder can refer to the description of Embodiment Four, and will not be repeated here.
[0068] In the embodiment of the present application, the N first deep spectral feature maps and the second deep spectral feature maps are subjected to adaptive fusion by the decoder to obtain the prediction map of the farmland in the farmland RGB image. The specific structure of the decoder can refer to the description of Embodiment Four, and will not be repeated here.
[0069] Example Six:
[0070] Figure 6 The structure of the farmland image segmentation device for fusing RGB and multispectral provided by the embodiment six of the present application is shown, only the parts related to the embodiment of the present application are shown for the convenience of description, which includes:
[0071] The first feature acquisition unit 61 is configured to perform feature extraction on the input farmland RGB image to obtain N first feature images, and perform feature enhancement on the N first feature images to obtain corresponding N first deep spectral feature maps, where N is a positive integer greater than 1.
[0072] The second feature acquisition unit 62 is configured to perform feature extraction on the input farmland multispectral image to obtain N second feature images, and perform feature enhancement on the N second feature images to obtain corresponding N second deep spectral feature maps, where the farmland multispectral image is a multispectral image of the farmland RGB image.
[0073] The feature fusion unit 63 is configured to perform adaptive fusion on the N first deep spectral feature maps and the N second deep spectral feature maps to obtain a prediction map of the farmland in the farmland RGB image.
[0074] In the embodiments of the present application, the units of the farmland image segmentation device can be realized by corresponding hardware or software units, and each unit can be an independent software or hardware unit, or can be integrated into a software or hardware unit, which does not limit the present application, and the specific implementation of each unit can refer to the description of the foregoing embodiments, which will not be repeated here.
[0075] Example Seven:
[0076] Figure 7 The structure of the computer device provided in the fourth embodiment of the present application is shown, and only the parts related to the embodiments of the present application are shown for the convenience of description.
[0077] The computer device 7 in the embodiments of the present application includes a processor 70, a memory 71, and a computer program 72 stored in the memory 71 and executable on the processor 70. The processor 70 implements the steps in the above-mentioned various farmland image segmentation method embodiments when executing the computer program 72, such as the steps S101-S103 shown in the above-mentioned method embodiments. Figure 1 Alternatively, the processor 70 implements the functions of the units in the above-mentioned various device embodiments when executing the computer program 72, such as the functions of the units 61-63 shown in the above-mentioned device embodiments. Figure 6 Alternatively, the processor 70 implements the functions of the units in the above-mentioned various device embodiments when executing the computer program 72, such as the functions of the units 61-63 shown in the above-mentioned device embodiments.
[0078] The computer device in the embodiments of the present application can be a personal computer or a server. The steps implemented when the processor 70 in the computer device 7 executes the computer program 72 to implement the farmland image segmentation method can refer to the description of the foregoing method embodiments, which will not be repeated here.
[0079] Example Five:
[0080] In the embodiments of the present application, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps in the above-mentioned farmland image segmentation method embodiments, for example, Figure 1 Or, the computer program is executed by the processor to implement the functions of the units in the above-mentioned various device embodiments, for example Figure 6 Or, the computer program is executed by the processor to implement the functions of the units in the above-mentioned various device embodiments, for example
[0081] The embodiments of the present application perform feature extraction on the input farmland RGB image to obtain N first feature images, perform feature enhancement on the N first feature images to obtain corresponding N first deep spectral feature maps, N is a positive integer greater than 1, perform feature extraction on the input farmland multispectral image to obtain N second feature images, perform feature enhancement on the N second feature images to obtain corresponding N second deep spectral feature maps, the farmland multispectral image is a multispectral image obtained by using a multispectral camera to shoot the same scene of the farmland RGB image, and perform adaptive fusion on the N first deep spectral feature maps and the second deep spectral feature maps to obtain a prediction map of the farmland in the farmland RGB image, thereby improving the segmentation accuracy of the farmland image in the farmland RGB image.
[0082] The computer readable storage medium of the embodiments of the present application can include any entity or device capable of carrying computer program code, recording medium, such as ROM / RAM, magnetic disk, optical disk, flash memory, etc.
[0083] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement and improvement within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for segmenting farmland images by fusing RGB and multispectral data, characterized in that, The method includes the following steps: Feature extraction is performed on the input RGB image of farmland to obtain N first feature images, and feature enhancement is performed on the N first feature images to obtain the corresponding N first deep spatial spectral feature maps, where N is a positive integer greater than 1; Feature extraction is performed on the input farmland multispectral image to obtain N second feature images. Feature enhancement is performed on the N second feature images to obtain corresponding N second deep spatial spectral feature maps. The farmland multispectral image is the multispectral image of the farmland RGB image. The N first deep spatial spectral feature maps and the second deep spatial spectral feature maps are adaptively fused by the decoder to obtain the predicted map of farmland in the farmland RGB image; Wherein, the value of N is 4, and the decoder includes a first, second, third, and fourth multimodal feature adaptive module, an upsampling layer, and a first, second, and third decoding module. Each multimodal feature adaptive module includes a first, second, third, fourth, fifth, and sixth convolutional layer and a normalized exponential function. The inputs of the first multimodal feature adaptive module are the first deep spatial-spectral feature map and the second deep spatial-spectral feature map. The output of the first multimodal feature adaptive module is the input of the upsampling layer. The inputs of the second multimodal feature adaptive module are the output of the upsampling layer, the first deep spatial-spectral feature map, and the second deep spatial-spectral feature map. The input of the first decoding module is the output of the second multimodal feature adaptive module. The inputs of the third multimodal feature adaptive module are the output of the first decoding module, the first deep spatial-spectral feature map, and the second deep spatial-spectral feature map. The input of the second decoding module is the output of the third multimodal feature adaptive module. The inputs of the fourth multimodal feature adaptive module are the output of the second decoding module, the first deep spatial-spectral feature map, and the second deep spatial-spectral feature map. The input of the third decoding module is the output of the fourth multimodal feature adaptive module. The output of the third decoding module is the prediction map.
2. The method as described in claim 1, characterized in that, The steps of extracting features from the input RGB image of farmland to obtain N first feature images, and then performing feature enhancement on the N first feature images, include: The input farmland RGB image is subjected to feature extraction by N convolutional pooling modules and one parallel convolutional module of the RGB image encoder to obtain the N first feature images, and the N first feature images are enhanced by N spatial spectral feature attention modules of the RGB image encoder.
3. The method as described in claim 2, characterized in that, The value of N is 4, and the RGB image encoder includes a first, a second, and a third convolutional pooling module, a first parallel convolution module, and a first, a second, a third, and a fourth spatial spectral feature attention module. The input to the first convolutional pooling module is the RGB image of the farmland. The output of the first convolutional pooling module is the input to the second convolutional pooling module and the first spatial spectral feature attention module. The output of the second convolutional pooling module is the input to the third convolutional pooling module and the second spatial spectral feature attention module. The output of the third convolutional pooling module is the input to the first parallel convolutional module and the third spatial spectral feature attention module. The output of the first parallel convolutional module is the input to the fourth spatial spectral feature attention module.
4. The method as described in claim 1, characterized in that, The steps of extracting features from the input multispectral image of farmland to obtain N second feature images, and enhancing the features of the N second feature images include: The input farmland multispectral image is used to extract features by N convolutional pooling modules and one parallel convolutional module of the multispectral image encoder to obtain N second feature images, and then the N second feature images are enhanced by N spatial spectral feature attention modules of the multispectral image encoder.
5. The method as described in claim 4, characterized in that, The value of N is 4, and the multispectral image encoder includes a fourth, a fifth, and a sixth convolutional pooling module, a second parallel convolution module, and a fifth, a sixth, a seventh, and an eighth spatial spectral feature attention module. The input to the fourth convolutional pooling module is the multispectral image of the farmland. The output of the fourth convolutional pooling module is the input to the fifth convolutional pooling module and the fifth spatial spectral feature attention module. The output of the fifth convolutional pooling module is the input to the sixth convolutional pooling module and the sixth spatial spectral feature attention module. The output of the sixth convolutional pooling module is the input to the second parallel convolutional module and the seventh spatial spectral feature attention module. The output of the second parallel convolutional module is the input to the eighth spatial spectral feature attention module.
6. The method as described in claim 1, characterized in that, The step of adaptively fusing the N first deep spatial spectral feature maps and the second deep spatial spectral feature maps includes: The first and second convolutional layers are used to perform convolution operations on the input first and second deep spatial spectral feature maps respectively to obtain the corresponding first and second intermediate feature maps. The first and second intermediate feature maps have the same number of channels. The third and fourth convolutional layers are used to perform convolution operations on the first and second intermediate feature maps respectively to compress the number of channels in the first and second intermediate feature maps, thereby obtaining the third and fourth intermediate feature maps. The third and fourth intermediate feature maps are spliced together to obtain the fifth intermediate feature map; The fifth convolutional layer is used to perform a convolution operation on the fifth intermediate feature map to compress the number of channels of the fifth intermediate feature map, thereby obtaining the sixth intermediate feature map; The first and second weight parameters corresponding to the sixth intermediate feature map are obtained through the normalized exponential function; Based on the first and second weight parameters and the first and second intermediate feature maps, the first and second weighted feature maps are obtained, and the first and second weighted feature maps are added together to obtain the seventh intermediate feature map. The sixth convolutional layer is used to perform a convolution operation on the seventh intermediate feature map to restore the number of channels in the seventh intermediate feature map, thereby obtaining a fused feature map.
7. A farmland image segmentation device integrating RGB and multispectral imaging, characterized in that, The device further includes: The first feature acquisition unit is used to extract features from the input farmland RGB image to obtain N first feature images, and to enhance the features of the N first feature images to obtain corresponding N first deep spatial spectral feature maps, where N is a positive integer greater than 1; The second feature acquisition unit is used to extract features from the input farmland multispectral image to obtain N second feature images, and to enhance the features of the N second feature images to obtain corresponding N second deep spatial spectral feature maps. The farmland multispectral image is a multispectral image of the farmland RGB image. The feature fusion unit is used to adaptively fuse the N first deep spatial spectral feature maps and the second deep spatial spectral feature maps through the decoder to obtain the predicted map of farmland in the farmland RGB image; Wherein, the value of N is 4, and the decoder includes a first, second, third, and fourth multimodal feature adaptive module, an upsampling layer, and a first, second, and third decoding module. Each multimodal feature adaptive module includes a first, second, third, fourth, fifth, and sixth convolutional layer and a normalized exponential function. The inputs of the first multimodal feature adaptive module are the first deep spatial-spectral feature map and the second deep spatial-spectral feature map. The output of the first multimodal feature adaptive module is the input of the upsampling layer. The inputs of the second multimodal feature adaptive module are the output of the upsampling layer, the first deep spatial-spectral feature map, and the second deep spatial-spectral feature map. The input of the first decoding module is the output of the second multimodal feature adaptive module. The inputs of the third multimodal feature adaptive module are the output of the first decoding module, the first deep spatial-spectral feature map, and the second deep spatial-spectral feature map. The input of the second decoding module is the output of the third multimodal feature adaptive module. The inputs of the fourth multimodal feature adaptive module are the output of the second decoding module, the first deep spatial-spectral feature map, and the second deep spatial-spectral feature map. The input of the third decoding module is the output of the fourth multimodal feature adaptive module. The output of the third decoding module is the prediction map.
8. The farmland image segmentation apparatus as described in claim 7, characterized in that, The first feature acquisition unit, when extracting features from the input RGB image of farmland to obtain N first feature images, and performing feature enhancement on the N first feature images, includes: The input farmland RGB image is subjected to feature extraction by N convolutional pooling modules and one parallel convolutional module of the RGB image encoder to obtain the N first feature images, and the N first feature images are enhanced by N spatial spectral feature attention modules of the RGB image encoder.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 6.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-encoder fused multispectral image semantic segmentation method
CN113762264A
Image segmentation method and device, electronic equipment and computer storage medium
CN115272354A