Methods, devices, media and equipment for detecting winter wheat under complex terrain conditions
By constructing a multimodal fusion segmentation network and combining optical and elevation features, the problem of accuracy in detecting winter wheat distribution in complex terrain was solved, and accurate detection of winter wheat distribution was achieved.
Patent Information
- Application Number
- CN202411360198.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-09-27
AI Technical Summary
Existing semantic segmentation methods have difficulty accurately identifying the spatial distribution of winter wheat in complex terrain, resulting in serious false alarms and missed detections, especially in hilly and mountainous areas where the distribution of winter wheat is relatively fragmented and irregular.
A multimodal training dataset, including dual-temporal optical remote sensing images and digital elevation model data, is used to construct a fusion segmentation network. Feature extraction and fusion are performed through the encoder and decoder. The multimodal feature fusion module and distributed feature enhancement module are used to improve the accuracy of the detection model.
It achieves accurate detection of winter wheat distribution in complex terrain areas, improves the effect of semantic segmentation, reduces false alarms and missed detections, and enhances detection accuracy.
Smart Images

Figure CN119339233B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image analysis technology, and in particular to a method, device, medium and equipment for detecting winter wheat under complex terrain conditions. Background Art
[0002] The development of satellite remote sensing technology has greatly facilitated the dynamic monitoring of global land surface changes, accumulating a vast amount of remote sensing satellite data that provides reliable evidence of land cover and utilization. In agriculture, remote sensing imagery can play a vital role in supporting rapid decision-making for resource management, agricultural production, and food security policymaking. Monitoring the growth and distribution of winter wheat, a key staple crop in my country, is particularly crucial for ensuring national food supply.
[0003] Existing techniques typically use semantic segmentation to classify surface objects at the pixel level using image processing techniques in remote sensing images of winter wheat to monitor their growth distribution. Current semantic segmentation methods, aside from traditional methods that simply distinguish ground objects based on edge or texture information and fixed spectral thresholds, are mostly based on the spectral information of single-phase optical images.
[0004] However, due to the influence of complex terrain, existing semantic segmentation methods are prone to lose the spatial distribution information of winter wheat when identifying winter wheat, which greatly reduces the accuracy of winter wheat distribution detection in complex terrain, resulting in serious false alarms and missed detections, especially in hilly and mountainous areas where winter wheat distribution is relatively fragmented and irregular, making it difficult to accurately capture small and scattered crop areas. Summary of the Invention
[0005] Based on this, it is necessary to provide a method, device, medium and equipment for detecting winter wheat under complex terrain conditions to address the above technical problems. This method can accurately detect the distribution of winter wheat in complex terrain areas.
[0006] The present invention adopts the following technical solutions:
[0007] Acquire a multimodal training dataset; the multimodal training dataset includes sample dual-temporal optical remote sensing images and sample digital elevation model data corresponding to a preset winter wheat area; the digital elevation model data is used to reflect the terrain information of the preset winter wheat area;
[0008] Based on the encoder and decoder, a fusion segmentation network for multimodal data is constructed. The encoder consists of two residual modules and four multimodal feature fusion modules with a dual-branch structure. The dual-branch structure includes an optical branch and an elevation branch. The decoder consists of four connected decoding modules, each of which includes an upsampling module and a distribution feature enhancement module. The inputs of the last three upsampling modules in the decoder are respectively connected to the outputs of the first three optical branches in the encoder.
[0009] The fusion segmentation network is trained using a multimodal training dataset to obtain a winter wheat detection model.
[0010] Obtain dual-temporal optical remote sensing images and digital elevation model data of the winter wheat area to be inspected;
[0011] The dual-phase optical remote sensing image and digital elevation model data are input into the winter wheat detection model, and two residual modules are used to unify the number of channels of the dual-phase optical remote sensing image and the digital elevation model data to obtain optical features and elevation features; the optical features and elevation features are sequentially extracted through four multimodal feature fusion modules, and the optical features and elevation features extracted by the fourth multimodal feature fusion module are spliced to obtain fused features; the fused features and the features output by the corresponding optical branch are used as the input of the decoder, and the fused features and the optical branches are feature enhanced through the decoding module in turn to obtain the detection results of the winter wheat area to be detected output by the decoder; the detection result is a pixel-level classification of the growth distribution of winter wheat in the winter wheat area to be detected.
[0012] Preferably, the acquiring of a multimodal training dataset includes:
[0013] Obtain sample optical remote sensing images and sample digital elevation model data of the preset winter wheat area at two time points;
[0014] The two sample optical remote sensing images are band-fused to obtain the sample dual-temporal optical remote sensing image;
[0015] The sample dual-temporal optical remote sensing images and sample digital elevation model data are used as training samples;
[0016] The training samples are divided into multimodal training dataset, multimodal test dataset and multimodal verification dataset according to the preset ratio.
[0017] Preferably, the method further comprises:
[0018] After obtaining the winter wheat detection model, the winter wheat detection model is tested using a validation data set to obtain test results;
[0019] If the test results do not meet the preset test conditions, the fusion segmentation network will continue to be trained using the multimodal training data set until a winter wheat detection model is obtained.
[0020] Preferably, the fusion segmentation network is trained using a multimodal training data set to obtain a winter wheat detection model, comprising:
[0021] Input the multimodal training dataset into the fusion segmentation network to obtain the predicted pixel-level classification of the multimodal training dataset;
[0022] The predicted pixel-level classification and the actual pixel-level classification corresponding to the multimodal training dataset are input into the cross entropy loss function to obtain the difference;
[0023] Iteratively update the fusion segmentation network based on difference, Adam optimizer and cosine annealing mechanism;
[0024] The iteratively updated fusion segmentation network was verified through the multimodal validation set, and the fusion segmentation network with the smallest loss function value on the multimodal validation set was determined as the winter wheat detection model.
[0025] Preferably, the optical branch of the multimodal feature fusion module consists of a maximum pooling submodule, a deformable convolutional network block and a residual network, and the elevation branch consists of a maximum pooling submodule and a convolutional block. The output of the convolutional block is connected to the output of the deformable convolutional network block along the channel direction as the input of the residual network.
[0026] The specific features for any multimodal feature fusion module include:
[0027] The optical features and elevation features are input into the maximum pooling submodule, and the maximum pooling operation is used to perform 2 times downsampling in the height and width directions to obtain the first pooled optical features and the first pooled elevation features;
[0028] The pooled elevation feature is extracted through the convolution block to obtain the second pooled elevation feature. The number of channels of the second pooled elevation feature is twice that of the first pooled elevation feature.
[0029] The first pooled optical feature is extracted through a deformable convolutional network block to obtain an intermediate optical feature; the number of channels of the intermediate optical feature is twice that of the first pooled optical feature;
[0030] The second pooled elevation features and the intermediate optical features are connected along the channel direction and input into the residual network to obtain the second pooled optical features; the second pooled optical features and the second pooled elevation features are the optical features and elevation features input into the next multimodal feature fusion module.
[0031] Preferably, the deformable convolutional network block includes a convolution layer, a batch normalization layer, a nonlinear rectification activation function layer and a variable convolution block; the convolution block consists of a standard 3×3 convolution with a padding of 1 and a stride of 1, a batch normalization layer and a nonlinear rectification activation function layer.
[0032] Preferably, the distribution feature enhancement module includes a global maximum pooling submodule, a global average pooling submodule, a first activation function, a convolutional layer, and a second activation function;
[0033] For any distribution feature enhancement module, it specifically includes:
[0034] The upsampled features are passed through the global maximum pooling submodule and the global average pooling submodule respectively, and the features after global maximum pooling and global average pooling are connected along the channel direction to obtain the intermediate weight matrix;
[0035] Pass the intermediate weight matrix through the convolution layer and the first activation function to obtain the spatial attention weight matrix;
[0036] The upsampled features are passed through the second activation function to obtain enhanced features;
[0037] The enhanced features, spatial attention weight matrix and sampled features are multiplied in turn to obtain enhanced features; the enhanced features are the input of the next decoding module.
[0038] The present invention provides a detection device for winter wheat under complex terrain conditions, comprising:
[0039] The training set acquisition module is used to obtain a multimodal training data set; the multimodal training data set includes sample dual-temporal optical remote sensing images and sample digital elevation model data corresponding to a preset winter wheat area; the digital elevation model data is used to reflect the terrain information of the preset winter wheat area;
[0040] A construction module is used to construct a fusion segmentation network for multimodal data based on an encoder and decoder. The encoder consists of two residual modules and four multimodal feature fusion modules with a dual-branch structure. The dual-branch structure includes an optical branch and an elevation branch. The decoder consists of four decoding modules connected together, each of which includes an upsampling module and a distribution feature enhancement module. The inputs of the last three upsampling modules in the decoder are respectively connected to the outputs of the first three optical branches in the encoder.
[0041] The training module is used to train the fusion segmentation network using a multimodal training dataset to obtain a winter wheat detection model;
[0042] A data acquisition module is used to obtain dual-temporal optical remote sensing images and digital elevation model data of the winter wheat area to be inspected;
[0043] The detection module is used to input the dual-phase optical remote sensing image and digital elevation model data into the winter wheat detection model, and use two residual modules to unify the number of channels of the dual-phase optical remote sensing image and the digital elevation model data to obtain optical features and elevation features; the optical features and elevation features are sequentially extracted through four multimodal feature fusion modules, and the feature map obtained by the fourth multimodal feature fusion module is spliced to obtain fusion features; the fusion features and the features output by the corresponding optical branch are used as inputs of the decoder, and the fusion features and the optical branches are sequentially enhanced through the decoding module to obtain the detection results of the winter wheat area to be detected output by the decoder; the detection result is a pixel-level classification of the growth distribution of winter wheat in the winter wheat area to be detected.
[0044] The present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it realizes the above-mentioned method for detecting winter wheat under complex terrain conditions.
[0045] The present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for detecting winter wheat under complex terrain conditions is implemented.
[0046] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects:
[0047] In the method for detecting winter wheat under complex terrain conditions provided by the present invention, optical features and elevation features are extracted and fused through a dual-branch structure of four multimodal feature fusion modules, so that the model can simultaneously utilize data from different sources, providing more comprehensive feature support for semantic segmentation; and, in the decoder, the outputs of the first three optical branches are connected to the inputs of the last three upsampling modules in the decoder, which can realize semantic reasoning of features at different levels. This combination effectively realizes the precise detection of complex terrain areas such as multi-scale and irregular crop areas, improves the effect of semantic segmentation, and thus improves the accuracy of winter wheat detection in complex areas.
[0048] Furthermore, the dual-phase optical remote sensing images and digital elevation model data respectively fuse the information obtained by different sensors, and make full use of the complementary information of the dual-phase and the terrain characteristics of the winter wheat area. The spatial distribution characteristics of the winter wheat area can be integrated into the model to provide a more comprehensive and accurate scene understanding, enabling the model to understand the research object more efficiently and deeply, thereby improving the accuracy of the winter wheat detection model and thus improving the accuracy of winter wheat detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0050] Figure 1 A schematic flow chart of a method for detecting winter wheat under complex terrain conditions provided by the present invention;
[0051] Figure 2 A schematic diagram of the structure of a fusion segmentation network provided by the present invention;
[0052] Figure 3 A schematic diagram of the structure of a multimodal feature fusion module provided by the present invention;
[0053] Figure 4 A structural diagram of a deformable convolutional network block provided by the present invention;
[0054] Figure 5 A schematic structural diagram of a distribution feature enhancement module provided by the present invention;
[0055] Figure 6 A schematic flow chart of another method for detecting winter wheat under complex terrain conditions provided by the present invention;
[0056] Figure 7 A schematic diagram of a winter wheat detection device under complex terrain conditions provided by the present invention;
[0057] Figure 8 A schematic diagram of a computer device for implementing a method for detecting winter wheat under complex terrain conditions provided by the present invention. DETAILED DESCRIPTION
[0058] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and corresponding drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0059] In order to solve the limitations of complex terrain and crop distribution detection, it is urgent to explore semantic segmentation methods based on multimodal data fusion. By combining multimodal data, the detection accuracy of crop distribution in complex terrain areas can be improved, the interference of factors such as lighting conditions, atmospheric conditions and terrain undulations on the classification results can be overcome, and false alarms and missed detections can be reduced. The problem of traditional crop segmentation methods losing spatial distribution information of winter wheat under complex terrain conditions and difficulty in accurately capturing small areas can be solved, thereby improving the applicability of winter wheat semantic segmentation models in complex terrain and improving the accuracy of winter wheat detection in complex terrain areas.
[0060] Based on this, the present invention provides a method for detecting winter wheat under complex terrain conditions. The technical solutions provided by various embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0061] The execution subject of the detection method of winter wheat under complex terrain conditions in the present invention can be a server, and the server can be a server set up on a business platform, or a device such as a desktop computer, a laptop computer, etc. that can execute the solution of the present invention.
[0062] Figure 1 The present invention is a flow chart of a method for detecting winter wheat under complex terrain conditions, which specifically includes the following steps:
[0063] S101, obtaining a multimodal training dataset.
[0064] The multimodal training dataset includes sample dual-temporal optical remote sensing images and sample digital elevation model data corresponding to a preset winter wheat area; the digital elevation model data is used to reflect the terrain information of the preset winter wheat area.
[0065] Preferably, obtaining a multimodal training data set includes: obtaining sample optical remote sensing images and sample digital elevation model data of a preset winter wheat area at two time points; and performing band fusion on the two sample optical remote sensing images to obtain sample dual-phase optical remote sensing images; and then using the sample dual-phase optical remote sensing images and the sample digital elevation model data as training samples; and finally dividing the training samples into a multimodal training data set, a multimodal test data set, and a multimodal verification data set according to a preset ratio.
[0066] The preset winter wheat region can be any region where winter wheat is grown; the two time points can be winter wheat and spring wheat, for example, December 2017 and March 2018. The sample optical remote sensing images can be optical remote sensing images collected in the preset winter wheat region at the two corresponding time points, and the sample digital elevation model images can be topographic information of the preset winter wheat region.
[0067] The two sample optical remote sensing images are subjected to band fusion to obtain a sample dual-temporal optical remote sensing image, specifically comprising: combining the two sample optical remote sensing images in the band dimension, stacking them in the order of the bands, and obtaining a sample dual-temporal optical remote sensing image containing a combination of two time bands; for example, the two sample optical remote sensing images are I1 and I2, each sample optical remote sensing image includes C bands, and the image size is H×W.
[0068] In order to make the multimodal input data containing dual-temporal information in the multimodal training dataset, first, I1 and I2 are band fused. Specifically, the spectral information of each pixel in I1 and I2 is cross-stacked to obtain a sample dual-temporal optical remote sensing image. The optical remote sensing image after band fusion changes from the original (C, H, W) to (2C, H, W). The image size of the sample dual-temporal optical remote sensing image is
[0069] Preferably, after obtaining the two sample optical remote sensing images by band fusion, graphic operations can be performed on the band-fused sample optical remote sensing images. Graphic operations balance data samples using image enhancement techniques, including operations such as flipping, rotating, scaling, translating, and cropping, to generate data samples of various perspectives and proportions, and to balance the number of data samples of rare types or sparsely distributed areas. The sample optical remote sensing images after image operations can be used as sample dual-temporal optical remote sensing images.
[0070] The sample digital elevation model (DEM) data is a spatial data model that describes the undulating morphological characteristics of the ground surface. It is a matrix composed of the elevation values of regular grid points on the ground, forming a grid structure dataset. As another data in the multimodal training dataset, it does not participate in the band fusion process, and the shape of the DEM data remains unchanged.
[0071] Therefore, the sample dual-time optical remote sensing images and sample digital elevation model data obtained above can be used as training samples.
[0072] Among them, the preset ratio can be 6:2:2, that is, the training samples are divided into a multimodal training data set, a multimodal test data set, and a multimodal verification data set according to the ratio of 6:2:2.
[0073] S102, constructing a fusion segmentation network for multimodal data based on the encoder and decoder; the encoder consists of two residual modules and four multimodal feature fusion modules with a dual-branch structure; the dual-branch structure includes an optical branch and an elevation branch; the decoder is composed of four decoding modules connected together, each decoding module includes an upsampling module and a distribution feature enhancement module; the inputs of the last three upsampling modules in the decoder are respectively connected to the outputs of the first three optical branches in the encoder.
[0074] like Figure 2 As shown, Figure 2 It can be a structural diagram of the fusion segmentation network; the fusion segmentation network consists of two parts: an encoder and a decoder. The encoder is responsible for extracting the features of multimodal data, and the decoder is used to restore the resolution and enhance the feature expression; among them, the multimodal data includes dual-phase optical remote sensing images and digital elevation model data.
[0075] The encoder consists of residual module 1, residual module 2, multimodal feature fusion module 1, multimodal feature fusion module 2, multimodal feature fusion module 3 and multimodal feature fusion module 4. The input of residual module 1 is digital elevation model data, and the input of residual module 2 is dual-temporal optical remote sensing image. Each multimodal feature fusion module (MFFM) has a dual-branch structure: an optical branch and an elevation branch, which extract the features of dual-temporal optical remote sensing image and digital elevation model data respectively. The features corresponding to dual-temporal optical remote sensing image can be called optical features, and the features corresponding to digital elevation model data can be called elevation features.
[0076] The multimodal feature fusion module 1 extracts features from the optical features and elevation features output by the residual module 1 and the residual module 2, respectively; the multimodal feature fusion module 2 extracts features from the optical features and elevation features output by the multimodal feature fusion module 1, respectively; the multimodal feature fusion module 3 extracts features from the optical features and elevation features output by the multimodal feature fusion module 2, respectively; the multimodal feature fusion module 4 extracts features from the optical features and elevation features output by the multimodal feature fusion module 3, respectively; finally, the optical features and elevation features output by the multimodal feature fusion module 4 can be feature fused. Specifically, the optical features and elevation features output by the multimodal feature fusion module 4 are connected by a concat operation, i.e., channel connection, to obtain fused features.
[0077] The multimodal feature fusion module combines optical imagery and digital elevation model data, extracting and fusing the features of both during the encoding process. The optical features of the optical imagery are extracted through deformable convolution (DCN) to capture the multi-scale and irregular crop distribution characteristics. Assuming that the optical feature is F opt , the elevation feature of the digital elevation model data is Fdem , then the feature fusion representation that can be obtained is: F fusion =concat(DCN(F opt ),Conv(F dem )).
[0078] The decoder is composed of four decoding modules: decoding module 1, decoding module 2, decoding module 3 and decoding module 4. Each decoding module consists of an upsampling module and a distribution feature enhancement module. The input of the upsampling module is the input of the decoding module to which it belongs, the output of the upsampling module is the input of the distribution feature enhancement module, and the output of the distribution feature enhancement module is the output of the decoding module to which it belongs; the input of decoding module 1 is the fusion feature output from the encoder, the input of decoding module 2 is the output of decoding module 1, the input of decoding module 3 is the output of decoding module 2, and the input of decoding module 4 is the output of decoding module 3.
[0079] Among them, the inputs of the last three upsampling modules in the decoder are respectively connected to the outputs of the first three optical branches in the encoder. Specifically, the output of the optical branch in the multimodal feature fusion module 1 is the input of the upsampling module in the decoding module 4, the output of the optical branch in the multimodal feature fusion module 2 is the input of the upsampling module in the decoding module 3, and the output of the optical branch in the multimodal feature fusion module 4 is the input of the upsampling module in the decoding module 2.
[0080] S103, training the fusion segmentation network using the multimodal training data set to obtain a winter wheat detection model.
[0081] Preferably, the fusion segmentation network is trained by a multimodal training data set to obtain a winter wheat detection model, including: inputting the multimodal training data set into the fusion segmentation network to obtain a predicted pixel-level classification of the multimodal training data set; inputting the predicted pixel-level classification and the true pixel-level classification corresponding to the multimodal training data set into the cross-entropy loss function to obtain a difference; based on the difference, Adam optimizer and cosine annealing mechanism, the fusion segmentation network is iteratively updated to obtain a winter wheat detection model.
[0082] The multimodal training dataset is input into the fusion segmentation network. The optical and elevation features of the crops are first captured by the multimodal feature fusion module in the encoder. Then, the response to the sparsely broken winter wheat area is further enhanced by the distribution feature enhancement module in the decoder stage. Finally, the enhanced features output by the decoder are obtained, which are the predicted pixel-level classifications of the multimodal training dataset.
[0083] During training, the fused segmentation network was optimized using a cross-entropy loss function to reduce model error. In each training iteration, the predicted pixel-level classification and the actual pixel-level classification corresponding to the multimodal training dataset were input into the cross-entropy loss function. The loss function calculated the prediction error for each pixel, i.e., the difference. The fused segmentation network was then iteratively updated based on the difference, the Adam optimizer, and a warm-up plus cosine annealing mechanism. Finally, the iteratively updated fused segmentation network was validated using a multimodal validation set. The fused segmentation network with the lowest loss function value on the multimodal validation set was selected as the winter wheat detection model.
[0084] The cross entropy loss function calculates the true pixel-level classification y for each pixel i and predict pixel-level classification The cross entropy loss function is shown in formula (1).
[0085]
[0086] Among them, L CE is the difference between the predicted pixel-level classification and the true pixel-level classification, and N is the number of pixels.
[0087] The fused segmentation network after each iterative update is verified through the multimodal validation set to obtain the loss function value of each iterative update. The fused segmentation network with the smallest loss function value on the multimodal validation set is determined as the winter wheat detection model.
[0088] Optionally, the fusion segmentation network is iteratively trained using the multimodal training dataset until the fusion segmentation network reaches a preset convergence condition, and the training is stopped. The convergence condition can be set to a number of iterations that meets a preset number of iterations.
[0089] In an exemplary embodiment, after obtaining a winter wheat detection model, a performance evaluation of the winter wheat can be performed. Specifically, the performance evaluation includes: after obtaining the winter wheat detection model, testing the winter wheat detection model using a validation dataset to obtain a test result; if the test result does not meet a preset test condition, continuing to train the fusion segmentation network using a multimodal training dataset until a winter wheat detection model is obtained. The preset test condition can be that the test accuracy reaches a preset accuracy threshold.
[0090] S104: Obtain dual-temporal optical remote sensing images and digital elevation model data of the winter wheat area to be inspected.
[0091] The winter wheat area to be detected may be an area where the winter wheat planting distribution needs to be detected, and the digital elevation model data is terrain information data of the winter wheat area to be detected.
[0092] Optical remote sensing images collected at two time points in the winter wheat area to be detected can be obtained first. The optical remote sensing images collected at the two time points can be wheat growth images of the winter wheat area to be detected in winter and wheat growth images in the spring of the following year.
[0093] The two optical remote sensing images are band-fused to obtain a dual-phase optical remote sensing image. It should be noted that the band fusion method is the same as the method of obtaining the sample dual-phase optical remote sensing image by band fusion in the above embodiment, and the embodiment of this application will not be repeated here.
[0094] S105, input the dual-phase optical remote sensing image and digital elevation model data into the winter wheat detection model, use two residual modules to unify the number of channels of the dual-phase optical remote sensing image and the digital elevation model data respectively, and obtain optical features and elevation features; extract the optical features and elevation features in turn through four multimodal feature fusion modules, and splice the optical features and elevation features obtained by feature extraction through the fourth multimodal feature fusion module to obtain fused features; use the fused features and the features output by the corresponding optical branch as input to the decoder, and enhance the features of the fused features and the optical branches in turn through the decoding modules to obtain the detection results of the winter wheat area to be detected output by the decoder.
[0095] The detection result is a pixel-level classification of the growth distribution of winter wheat in the winter wheat area to be detected. The pixel-level classification can be a pixel-level segmentation map showing the distribution of winter wheat in the winter wheat area to be detected.
[0096] Specifically, the dual-phase optical remote sensing image and digital elevation model data are input into the winter wheat detection model. First, the number of channels of the dual-phase optical remote sensing image and digital elevation model data are changed to 64 respectively through the residual module, while the size remains unchanged; then the features are extracted through the multimodal feature fusion module.
[0097] like Figure 3 As shown, Figure 3 The schematic diagram of the structure of the multimodal feature fusion module is shown in Figure 2. The optical branch of the multimodal feature fusion module consists of a maximum pooling submodule, a deformable convolutional network block, and a residual network. The elevation branch consists of a maximum pooling submodule and a convolutional block. The output of the convolutional block is connected to the output of the deformable convolutional network block along the channel direction as the input of the residual network. Taking the i-th layer multimodal feature fusion module as an example, the input optical feature is The elevation feature is The output optical characteristics are The elevation feature is The convolutional block consists of a standard 3×3 convolution with padding of 1 and stride of 1, a batch normalization layer, and a nonlinear rectification activation function layer.
[0098] The specific steps for any multimodal feature fusion module include: inputting the optical features and elevation features into the maximum pooling submodule, performing 2-fold downsampling in the height and width directions through the maximum pooling operation to obtain the first pooled optical features and the first pooled elevation features; extracting the pooled elevation features through the convolution block to obtain the second pooled elevation features, and the number of channels of the second pooled elevation features is twice that of the first pooled elevation features; then extracting the first pooled optical features through the deformable convolution network block to obtain the intermediate optical features; the number of channels of the intermediate optical features is twice that of the first pooled optical features; finally, connecting the second pooled elevation features and the intermediate optical features along the channel direction and inputting them into the residual network to obtain the second pooled optical features; the second pooled optical features and the second pooled elevation features are the optical features and elevation features input to the next multimodal feature fusion module.
[0099] Taking the i-th layer multimodal feature fusion module as an example, assuming that the elevation branch input of the i-th layer multimodal feature fusion module is a feature map of shape (C, H, W) After the maximum pooling operation (maxpool) is performed with 2 times downsampling in width and height, it becomes the first pooled elevation feature of (C, H / 2, W / 2), which is then sent to the convolution block (Conv block) to extract features, and the second pooled elevation feature is output through the convolution block module. The number of channels is doubled and the size is (2C, H / 2, W / 2). The multimodal feature fusion module is combined with the residual module to maintain feature consistency. The output of the elevation branch is:
[0100]
[0101] The maxpool() in the formula has a pooling kernel of 2 and a stride of 2; the Convblock block consists of a standard 3×3 convolution with a padding of 1 and a stride of 1, a batch normalization (BN) layer, and a nonlinear rectification (ReLU) activation function layer.
[0102] The optical branch consists of a maximum pooling submodule, a deformable convolutional network block (DCN block) and a residual network. The deformable convolutional network block includes a convolution layer, a batch normalization layer, a nonlinear rectification activation function layer and a variable convolution block. Specifically, Figure 4 As shown, Figure 4The following is a structural diagram of the deformable convolutional network block. The deformable convolutional network block includes a 3×3 convolution layer, a batch normalization (BN) layer, a nonlinear rectified linear unit (ReLU) activation function, a deformable convolutional network version 2 (DCNv2) block, a batch normalization layer, and a nonlinear activation function. For example, the input of the deformable convolutional network block is The output is
[0103] Taking the i-th layer multimodal feature fusion module as an example, let the input feature map of the optical branch of the i-th layer multimodal feature fusion module be The maximum pooling operation is used to perform 2x downsampling in the width and height directions to obtain the first pooled optical features. Then, the first pooled optical features are extracted in the deformable convolutional network block to extract the remote sensing image features, and the number of channels in the remote sensing image feature map is doubled to obtain the intermediate optical features. and the second pooled elevation feature from the elevation branch After connecting along the channel direction, it is sent into the residual structure, and the final output is the second pooled optical feature The calculation formula for the optical branch is as follows:
[0104]
[0105] The calculation formulas of DCNblock and Resblock are as follows; among them, DCNv2 is deformable convolution.
[0106] DCNblock(x)=ReLU(BN(DCNv2(x))) (4)
[0107] Resblock(x)=x+Convblock(Convblock(x)) (5)
[0108] For any upsampling module, the upsampling module performs a 2x upsampling of the input feature map through bilinear interpolation to restore the resolution of the input image. It should be noted that the inputs of upsampling modules 2, 3, and 4 are the output of the previous distribution feature enhancement module and the output of the corresponding optical branch. Therefore, before upsampling, the upsampling module can first connect the two input features to obtain a feature map, and then upsample the feature map.
[0109] The Distribution Feature Enhancement Module (DFEM) uses the weighted attention mechanism in the decoding stage to enhance the feature expression of the winter wheat distribution area. Figure 5 As shown, the distribution feature enhancement module includes a global maximum pooling submodule, a global average pooling submodule, a first activation function, a convolutional layer and a second activation function.
[0110] For any distribution feature enhancement module, it specifically includes: passing the upsampled features through the global maximum pooling submodule and the global average pooling submodule respectively, and connecting the features after global maximum pooling and global average pooling along the channel direction to obtain an intermediate weight matrix; passing the intermediate weight matrix through the convolution layer and the first activation function to obtain the spatial attention weight matrix; passing the upsampled features through the second activation function to obtain enhanced features; performing point multiplication on the enhanced features, the spatial attention weight matrix and the sampled features in turn to obtain enhanced features; the enhanced features are the input of the next decoding module.
[0111] Specifically, suppose the feature map x of the i-th layer input i The size of is (C, H, W), and after upsampling, it becomes (C, 2H, 2W), and then it is concatenated with the feature map of the same size as the decoder, and the output feature map size becomes (2C, 2H, 2W).
[0112] The features after sampling above are x i For example, the upsampled features are respectively passed through global max pooling and global mean pooling, and the features after global max pooling and global mean pooling are connected along the channel direction to obtain the intermediate weight matrix; then the intermediate weight matrix passes through a 3×3 convolution layer and the first activation function sigmold to obtain the spatial attention weight matrix W i sam , spatial attention weight matrix W i sam Used to enhance wheat distribution information.
[0113] W i sam =sigmoid(Conv(concat(globalmaxpool(x i ),globalavgpool(x i ))))(6)
[0114] The distribution feature enhancement module converts x i The feature strength is adjusted by the second activation function tanh()+1, and the enhanced features generated by the second activation function tanh()+1 are and the spatial attention weight matrix Perform point multiplication to generate the distribution feature enhancement weight matrix W i DFEM , strengthen the wheat area, and then W i DFEM With the upsampled feature x i Perform dot multiplication to obtain the enhanced feature x i+1 The calculation process is as follows:
[0115]
[0116] x i+1 =W i DFEM ×x i (9)
[0117] It should be noted that the enhanced features output by the distribution feature enhancement module 4 are the pixel-level classifications of the winter wheat region to be detected output by the decoder.
[0118] In an exemplary embodiment, a fusion segmentation architecture design based on multimodal data integrates data characteristics of different modalities by constructing a multimodal feature fusion module and a distribution feature enhancement module to enhance the feature expression capability of the crop area.
[0119] Winter wheat detection based on multimodal fusion segmentation network is mainly divided into three parts: input, training and output; Figure 6 As shown, Figure 6 A flow chart of a method for detecting winter wheat under complex terrain conditions, the method comprising:
[0120] S601, multimodal dataset construction.
[0121] Obtain fused dual-temporal optical remote sensing images and digital elevation model data, and divide them into multimodal training dataset, multimodal test dataset, and multimodal validation dataset in a ratio of 6:2:2.
[0122] S602, multimodal feature fusion.
[0123] S603, feature enhancement and restoration.
[0124] Based on the encoder and decoder, a fusion segmentation network for multimodal data is constructed. The encoder consists of two residual modules and four multimodal feature fusion modules with a dual-branch structure. The dual-branch structure includes an optical branch and an elevation branch. The decoder consists of four connected decoding modules, each of which includes an upsampling module and a distribution feature enhancement module. The inputs of the last three upsampling modules in the decoder are connected to the outputs of the first three optical branches in the encoder. Figure 2-Figure 5 .
[0125] S604, model training and loss optimization.
[0126] The multimodal training dataset is input into the fusion segmentation network, first passing through the multimodal feature fusion module in the encoder to capture the crop's spectral characteristics (DCN) and terrain elevation information (DEM). Then, in the decoder stage, the distribution feature enhancement module further enhances the response to sparsely broken winter wheat areas. During training, a cross-entropy loss function is used for optimization, gradually reducing the model's error. In each training iteration, the loss function calculates the prediction error for each pixel and continuously updates the model's parameters through backpropagation to optimize segmentation performance. As training progresses, the model's detection accuracy for winter wheat areas gradually improves and the error gradually decreases, ensuring that even small areas in complex terrain are accurately captured.
[0127] S605, result generation and performance evaluation.
[0128] The model with the lowest loss on the multimodal validation dataset is saved as the best model, which is then tested on the multimodal test dataset. After passing through the decoder output layer, the winter wheat interval results are generated. This result is a pixel-level segmentation map showing the distribution of winter wheat in complex terrain.
[0129] This application fully considers the spatial distribution characteristics of winter wheat under complex terrain, utilizes the complementary advantages of multimodal data, and solves the problem that traditional crop segmentation methods are difficult to capture small crop areas and lose spatial distribution information in complex terrain by constructing a multimodal feature fusion module and a distribution feature enhancement module. It effectively realizes the precise detection of multi-scale and irregular crop areas, and ultimately significantly improves the accuracy and stability of winter wheat distribution detection under complex terrain conditions.
[0130] In the winter wheat detection method under complex terrain conditions provided by the present invention, a fusion segmentation structure is designed, and the dual-phase optical remote sensing image and digital elevation model data are used as a multimodal training data set to train the fusion segmentation structure to obtain a winter wheat detection model, so as to determine the detection results of the winter wheat area to be identified through the winter wheat detection model.
[0131] The fusion segmentation structure includes an encoder and a decoder. In the encoder, two residual modules can effectively extract the features of dual-temporal optical remote sensing images and digital elevation model data, enhancing the model's expressiveness. The dual-branch structure of the multimodal feature fusion module effectively fuses optical and elevation features, enabling the model to simultaneously utilize data from different sources.
[0132] Moreover, in the decoder, the upsampling module and the distribution feature enhancement module can better restore the spatial resolution and enhance the feature expression ability. At the same time, the output of the optical branch is connected to the input of the upsampling module of the decoder, so that the model can fuse data at different levels and utilize the features of each level to make the feature transfer more comprehensive, effectively realizing the accurate recognition of multi-scale and irregular crop areas, thereby improving the accuracy of winter wheat detection results in complex terrain areas.
[0133] In addition, since the dual-phase optical remote sensing images and digital elevation model data respectively integrate the information obtained by different sensors and make full use of the complementary information of the dual temporal phases, they can provide a more comprehensive and accurate scene understanding, enabling the model to understand the research object more efficiently and deeply, thereby improving the accuracy of the winter wheat detection model and thus improving the accuracy of the winter wheat detection results.
[0134] The above is a method for detecting winter wheat under complex terrain conditions provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding device for detecting winter wheat under complex terrain conditions, such as Figure 7 shown.
[0135] Figure 7 This is a schematic diagram of a device for detecting winter wheat under complex terrain conditions provided by the present invention. The device 700 includes:
[0136] The training set acquisition module 701 is used to acquire a multimodal training data set; the multimodal training data set includes sample dual-temporal optical remote sensing images and sample digital elevation model data corresponding to a preset winter wheat area; the digital elevation model data is used to reflect the terrain information of the preset winter wheat area;
[0137] A construction module 702 is configured to construct a fusion segmentation network for multimodal data based on an encoder and a decoder; the encoder is composed of two residual modules and four multimodal feature fusion modules with a dual-branch structure; the dual-branch structure includes an optical branch and an elevation branch; the decoder is composed of four decoding modules connected together, each of which includes an upsampling module and a distribution feature enhancement module; the inputs of the last three upsampling modules in the decoder are respectively connected to the outputs of the first three optical branches in the encoder;
[0138] A training module 703 is used to train the fusion segmentation network using a multimodal training data set to obtain a winter wheat detection model;
[0139] The data acquisition module 704 is used to acquire dual-temporal optical remote sensing images and digital elevation model data of the winter wheat area to be detected;
[0140] Detection module 705 is used to input the dual-phase optical remote sensing image and digital elevation model data into the winter wheat detection model, use two residual modules to unify the number of channels of the dual-phase optical remote sensing image and the digital elevation model data respectively, and obtain optical features and elevation features; extract the optical features and elevation features in turn through four multimodal feature fusion modules, and splice the feature map obtained by the fourth multimodal feature fusion module to obtain fusion features; use the fusion features and the features output by the corresponding optical branch as input to the decoder, and enhance the features of the fusion features and the optical branches in turn through the decoding modules to obtain the detection result of the winter wheat area to be detected output by the decoder; wherein the detection result is a pixel-level classification of the growth distribution of winter wheat in the winter wheat area to be detected.
[0141] For the specific limitations of the detection device for winter wheat under complex terrain conditions, please refer to the limitations of the detection method for winter wheat under complex terrain conditions above, and will not be repeated here. The various modules in the above-mentioned detection device for winter wheat under complex terrain conditions can be implemented in whole or in part through software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0142] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 Provided is a detection method for winter wheat under complex terrain conditions.
[0143] The present invention also provides Figure 8 The structural diagram of the computer equipment shown in FIG. Figure 8 As shown in the figure, at the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Provided is a detection method for winter wheat under complex terrain conditions.
[0144] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0145] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present invention.
Claims
1. A method for detecting winter wheat under complex terrain conditions, characterized in that: include: Obtain a multimodal training dataset; The multimodal training dataset includes sample dual-temporal optical remote sensing images and sample digital elevation model data corresponding to the preset winter wheat area; Digital elevation model data is used to reflect the terrain information of the preset winter wheat area; Based on the encoder and decoder, a fusion segmentation network for multimodal data is constructed; the encoder consists of two residual modules and four multimodal feature fusion modules with a dual-branch structure; the dual-branch structure includes an optical branch and an elevation branch; The decoder consists of four connected decoding modules, each of which includes an upsampling module and a distribution feature enhancement module. The inputs of the last three upsampling modules in the decoder are connected to the outputs of the first three optical branches in the encoder respectively. The fusion segmentation network is trained using a multimodal training dataset to obtain a winter wheat detection model. Obtain dual-temporal optical remote sensing images and digital elevation model data of the winter wheat area to be inspected; The dual-phase optical remote sensing image and digital elevation model data are input into the winter wheat detection model, and two residual modules are used to unify the number of channels of the dual-phase optical remote sensing image and the digital elevation model data to obtain optical features and elevation features; the optical features and elevation features are sequentially extracted through four multimodal feature fusion modules, and the optical features and elevation features extracted by the fourth multimodal feature fusion module are spliced to obtain fused features; the fused features and the features output by the corresponding optical branch are used as input to the decoder, and the fused features and the optical branches are sequentially enhanced through the decoding module to obtain the detection results of the winter wheat area to be detected output by the decoder; the detection results are pixel-level classification of the growth distribution of winter wheat in the winter wheat area to be detected.
2. The method according to claim 1, characterized in that The obtaining of a multimodal training data set includes: Obtain sample optical remote sensing images and sample digital elevation model data of the preset winter wheat area at two time points; The two sample optical remote sensing images are band-fused to obtain the sample dual-temporal optical remote sensing image; The sample dual-temporal optical remote sensing images and sample digital elevation model data are used as training samples; The training samples are divided into multimodal training dataset, multimodal test dataset and multimodal verification dataset according to the preset ratio.
3. The method according to claim 2, characterized in that The method further comprises: After obtaining the winter wheat detection model, the winter wheat detection model is tested using a validation data set to obtain test results; If the test results do not meet the preset test conditions, the fusion segmentation network will continue to be trained using the multimodal training data set until a winter wheat detection model is obtained.
4. The method according to claim 1, wherein The fusion segmentation network is trained using a multimodal training data set to obtain a winter wheat detection model, including: Input the multimodal training dataset into the fusion segmentation network to obtain the predicted pixel-level classification of the multimodal dataset; The predicted pixel-level classification and the actual pixel-level classification corresponding to the multimodal training dataset are input into the cross entropy loss function to obtain the difference; Iteratively update the fusion segmentation network based on difference, Adam optimizer and cosine annealing mechanism; The iteratively updated fusion segmentation network was verified through the multimodal validation set, and the fusion segmentation network with the smallest loss function value on the multimodal validation set was determined as the winter wheat detection model.
5. The method according to claim 1, wherein The optical branch of the multimodal feature fusion module consists of a maximum pooling submodule, a deformable convolutional network block and a residual network, and the elevation branch consists of a maximum pooling submodule and a convolutional block. The output of the convolutional block is connected to the output of the deformable convolutional network block along the channel direction as the input of the residual network. The specific features of any multimodal feature fusion module include: The optical features and elevation features are input into the maximum pooling submodule, and the maximum pooling operation is used to perform 2 times downsampling in the height and width directions to obtain the first pooled optical features and the first pooled elevation features; The pooled elevation feature is extracted through the convolution block to obtain the second pooled elevation feature. The number of channels of the second pooled elevation feature is twice that of the first pooled elevation feature. The first pooled optical feature is extracted through a deformable convolutional network block to obtain an intermediate optical feature; the number of channels of the intermediate optical feature is twice that of the first pooled optical feature; The second pooled elevation features and the intermediate optical features are connected along the channel direction and input into the residual network to obtain the second pooled optical features; the second pooled optical features and the second pooled elevation features are the optical features and elevation features input into the next multimodal feature fusion module.
6. The method according to claim 5, characterized in that The deformable convolutional network block includes a convolution layer, a batch normalization layer, a nonlinear rectification activation function layer, and a variable convolution block; the convolution block consists of a standard 3×3 convolution with a padding of 1 and a stride of 1, a batch normalization layer, and a nonlinear rectification activation function layer.
7. The method according to claim 1, characterized in that The distribution feature enhancement module includes a global maximum pooling submodule, a global average pooling submodule, a first activation function, a convolutional layer, and a second activation function; For any distribution feature enhancement module, it specifically includes: The upsampled features are passed through the global maximum pooling submodule and the global average pooling submodule respectively, and the features after global maximum pooling and global average pooling are connected along the channel direction to obtain the intermediate weight matrix; Pass the intermediate weight matrix through the convolution layer and the first activation function to obtain the spatial attention weight matrix; The upsampled features are passed through the second activation function to obtain enhanced features; The enhanced features, spatial attention weight matrix and sampled features are multiplied in turn to obtain enhanced features; the enhanced features are the input of the next decoding module.
8. A detection device for winter wheat under complex terrain conditions, characterized in that: include: The training set acquisition module is used to obtain multimodal training data sets; The multimodal training dataset includes sample dual-temporal optical remote sensing images and sample digital elevation model data corresponding to the preset winter wheat area; Digital elevation model data is used to reflect the terrain information of the preset winter wheat area; A construction module is used to construct a fusion segmentation network for multimodal data based on an encoder and decoder; the encoder consists of two residual modules and four multimodal feature fusion modules with a dual-branch structure; the dual-branch structure includes an optical branch and an elevation branch; The decoder consists of four connected decoding modules, each of which includes an upsampling module and a distribution feature enhancement module. The inputs of the last three upsampling modules in the decoder are connected to the outputs of the first three optical branches in the encoder respectively. The training module is used to train the fusion segmentation network using a multimodal training dataset to obtain a winter wheat detection model; A data acquisition module is used to obtain dual-temporal optical remote sensing images and digital elevation model data of the winter wheat area to be inspected; The detection module is used to input dual-phase optical remote sensing images and digital elevation model data into the winter wheat detection model, use two residual modules to unify the number of channels of the dual-phase optical remote sensing images and digital elevation model data respectively, and obtain optical features and elevation features; extract the optical features and elevation features in turn through four multimodal feature fusion modules, and splice the feature map obtained by the fourth multimodal feature fusion module to obtain fusion features; use the fusion features and the features output by the corresponding optical branches as inputs of the decoder, and enhance the features of the fusion features and the optical branches in turn through the decoding modules to obtain the detection results of the winter wheat area to be detected output by the decoder; the detection results are pixel-level classification of the growth distribution of winter wheat in the winter wheat area to be detected.
9. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and capable of running on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Remote sensing image building extraction method based on multi-scale feature fusion and enhancement
CN114387512A
Winter wheat yield prediction method, device and equipment based on multi-modal canopy image
CN116307105A