Method for extracting lodging winter wheat spatial distribution based on bimodal feature fusion
By constructing a dual-stream semantic segmentation network DMFCFNet, the feature correction and fusion of remote sensing images and DSM elevation images is solved, and the problem of low automation degree of spatial distribution extraction of lodged winter wheat in drone remote sensing images is achieved, achieving higher extraction accuracy and boundary clarity.
Patent Information
- Application Number
- CN202510702843.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The prior art when extracting the spatial distribution of lodged winter wheat from drone remote sensing images, the degree of automation is low, it is easily disturbed by environmental factors, and the interaction efficiency of cross-modal features is low, resulting in poor recognition performance.
A method of spatial distribution extraction of lodged winter wheat with bimodal feature fusion is proposed. By constructing a dual-flow semantic segmentation network DMFCFNet, the parallel input of remote sensing images and DSM elevation images is used to perform feature correction and fusion, dynamically adjust the receptive field range, and enhance the discrimination ability of features.
The extraction accuracy of the lodged winter wheat area is improved, the loss of small-scale features in the deep network is reduced, and the regional continuity and boundary clarity of segmentation results are improved.
Smart Images

Figure CN120236201A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of agricultural remote sensing monitoring, and particularly relates to a method for extracting the spatial distribution of lodging winter wheat by dual-modal feature fusion. Background Art
[0002] Winter wheat is one of the main food crops in China, and its stable and high yield is of great strategic significance for national food security and optimizing the allocation of agricultural resources. In actual production, the lodging of winter wheat is the main restricting factor for its high yield, stable yield and high quality. How to quickly and accurately obtain the spatial distribution information data of lodging winter wheat by using unmanned aerial vehicle (UAV) remote sensing images has become an issue that agricultural workers are extremely concerned about.
[0003] UAV remote sensing, with its sub-meter spatial resolution and flexible acquisition advantages, can make up for the disadvantages of satellite remote sensing in spatio-temporal resolution and provide an important data basis for the field of precision agriculture research. Early methods for extracting the spatial distribution of lodging winter wheat from UAV images mainly relied on features such as color, texture and shape, through image processing algorithms, such as threshold-based segmentation methods, watershed algorithms and edge detection methods, etc. However, these algorithms rely on manually selected features, have low automation, are easily interfered by environmental factors such as texture and terrain, and have certain limitations. With the development of machine learning, researchers have applied algorithms such as support vector machines, random forests and decision trees to the field of UAV remote sensing image segmentation. By learning existing features from the data and constructing a classification model based on the training data, the accuracy of image segmentation has been improved. However, with the progress of remote sensing technology, the semantic information of images has become more abundant, resulting in more prominent problems of same object with different spectra and inter-class scale imbalance, making it prone to mis-segmentation phenomena when dealing with complex terrain or regions with similar features.
[0004] In recent years, convolutional neural networks have made breakthrough progress in the field of remote sensing image processing with their excellent feature learning and extraction capabilities. Through a multi-level feature abstraction mechanism, they can more effectively extract key semantic features such as texture, spectrum and geometry in remote sensing images, and improve the accuracy of per-pixel segmentation of images. However, when the existing methods are used to extract the lodging area of winter wheat in remote sensing images, they still face problems such as small area of lodging targets, discontinuous distribution and significant scale changes. They fail to fully consider the non-linear correlation features between different modalities in the lodging area, resulting in low cross-modal feature interaction efficiency and affecting the recognition performance of the lodging area. Moreover, currently, multi-scale feature extraction networks are difficult to dynamically adjust the receptive field range, and there is a problem of insufficient integration of local detail features and global context information, resulting in the gradual dilution of small-scale lodging area features in the deep network, making the segmentation results prone to phenomena such as regional fracture, blurred boundary and missed detection of small targets. Summary of the Invention
[0005] To solve the above technical problems, the present application proposes the following technical solutions: In a first aspect, an embodiment of the present application provides a method for extracting the spatial distribution of lodging winter wheat with dual-modal feature fusion, including: Establish a dual-stream semantic segmentation network DMFCFNet and train DMFCFNet using the constructed training sample dataset, and obtain the optimal DMFCFNet after testing with the test sample dataset; Parallelly input the preprocessed remote sensing image and DSM elevation image into DMFCFNet to obtain the semantic texture feature information of the remote sensing image and the spatial structure feature information of the DSM elevation image at different scales; Perform feature correction and fusion on the semantic texture feature information of the remote sensing image and the spatial structure feature information of the DSM elevation image to obtain the fused feature information; Perform segmentation prediction on the fused feature information at multiple different scales to achieve the extraction of the spatial distribution information of the winter wheat lodging area.
[0006] In a possible implementation manner, the DMFCFNet includes: an encoder and a decoder. The encoder includes a remote sensing image branch and a DSM elevation image branch. The remote sensing image branch and the DSM elevation image branch include multiple stage residual modules and multiple cross-module feature correction modules. Each stage residual module is correspondingly connected to a cross-module feature correction module, and the output of the cross-module feature correction module is connected to the adaptive dual-modal feature fusion module; The remote sensing image and the DSM elevation image are parallelly input into the residual modules of the remote sensing image branch and the DSM elevation image branch, and feature alignment and complementarity are achieved through the cross-module feature correction module; The adaptive dual-modal feature fusion module adaptively fuses the corrected remote sensing image features and DSM elevation image features; The decoder includes a multi-scale aggregation decoding module and a classification module. The multi-scale aggregation decoding module is connected to each adaptive dual-modal feature fusion module in the encoder for receiving the fused features output by different adaptive dual-modal feature fusion modules. The multi-scale aggregation decoding module is connected to the classification module. The decoder is used to perform segmentation prediction on the fused feature information at different scales to achieve classification.
[0007] In a possible implementation manner, the establishment of the dual-stream semantic segmentation network DMFCFNet and the training of DMFCFNet using the constructed training sample dataset include: Determine the hyperparameters during the training process and initialize each parameter in the network; Input the selected training image - DSM data sample - label pairs as training sample data into DMFCFNet; Use DMFCFNet to perform a forward calculation on the current training data and calculate the loss according to the Loss loss function; Use the SGD algorithm to update the parameters of DMFCFNet until the loss function is less than the specified expected value or the loss value no longer changes.
[0008] In a possible implementation, the calculation formula of the Loss loss function is: where, is the cross - entropy loss, is the Dice loss, and are hyperparameters, n is the number of samples, is the label value of the i - th sample, is the predicted value of the model for the i - th sample, and N is the total number of pixel points.
[0009] In a possible implementation, the parallel input of the pre - processed remote sensing image and DSM elevation image into DMFCFNet to obtain the semantic texture feature information of the remote sensing image and the spatial structure feature information of the DSM elevation image includes: After parallel input of the pre - processed remote sensing image and DSM elevation image into DMFCFNet, the input layer of the remote sensing image branch uses large - kernel convolution to construct a wide receptive field, capture the low - frequency texture features in the input remote sensing image, and downsample to reduce the size of the remote sensing image; The DSM elevation image branch adjusts the number of convolutional kernel channels of the input layer for single - channel elevation data, reveals the spatial correlation between terrain undulation and crop lodging through elevation gradient features, and downsamples to reduce the size of the DSM elevation image; The first - stage residual module retains the high - resolution details of the remote sensing image and DSM elevation image through skip connections, compresses the feature map with max - pooling at the end, and extracts deep semantic features through extended residual units.
[0010] In a possible implementation, perform feature correction and fusion on the semantic texture feature information of the remote sensing image and the spatial structure feature information of the DSM elevation image to obtain the fused feature information, including: Global average pooling and max pooling are respectively performed on the input remote sensing image and DSM elevation image. After capturing the mean information and extreme value responses at the channel level, a non-linear transformation is performed through a convolutional module to generate channel attention weights. The semantic clues of the two poolings are fused by element-wise addition, and finally, an adaptive weighting in the channel dimension is achieved through the sigmoid activation function, and the output generates the corrected features. and ; The corrected features are processed using multi-scale depthwise separable convolutions with standard convolutions and asymmetric convolution kernels. and ; The corrected features and are dynamically complemented through a cross-modal cross-attention module. The adaptive dual-modal feature fusion module realizes the adaptive fusion of the remote sensing image and DSM elevation image features through a structured cross-modal interaction and dynamic spatial fusion mechanism.
[0011] In one possible implementation, the dynamic complementation of the corrected features and through the cross-modal cross-attention module includes: Stitch the dual-modal features of the remote sensing image and DSM elevation image. Extract cross-modal associations through 3×3 convolution and RELU activation function, and generate spatial weights by 1×1 convolution. and ; Use two learnable parameters and to perform weighted fusion on the dual-modal features to generate output features. The calculation formula is: where , respectively represent the remote sensing image and DSM elevation image, and represent learnable parameters, , represent the corrected remote sensing image features and corrected DSM elevation image features, and represent spatial weights, represents element-wise multiplication.
[0012] In a possible implementation, the adaptive bimodal feature fusion module realizes the adaptive fusion of remote sensing image and DSM elevation image features through a structured cross-modal interaction and dynamic spatial fusion mechanism, including: The adaptive bimodal feature fusion module adopts a dual-branch parallel architecture to separately process the corrected remote sensing image features and DSM elevation image features; Each branch uses an independent convolutional module to expand the input channels and divides them into three groups of feature maps along the channel dimension: guiding features, associated features, and content features; The guiding features of the remote sensing image are concatenated with the associated features of the DSM elevation image and the content features of the DSM elevation image along the channel dimension to form a first fusion feature; At the same time, the guiding features of the DSM elevation image are reversely concatenated with the associated features of the remote sensing image and the content features of the remote sensing image to form a second fusion feature; The first fusion feature and the second fusion feature are respectively compressed to the hidden channel dimension through convolution to generate a first compact representation and a second compact representation; The first compact representation and the second compact representation are mapped back to the original number of channels through a projection layer, and are respectively added pixel by pixel to the input corrected remote sensing image features and DSM elevation image features; A spatial dynamic weight map is generated through a spatial adaptive feature fusion network to realize the adaptive fusion of remote sensing image and DSM elevation image features.
[0013] In a possible implementation, the extraction of the spatial distribution information of the winter wheat lodging area is realized by segmenting and predicting the fused feature information at multiple different scales, including: The fused feature information at multiple different scales is aligned in the channel dimension, and a parallel convolutional layer is used to uniformly map the features of each layer; After eliminating the cross-modal feature distribution differences through group normalization and GELU activation function, a top-down information fusion path is adopted. Starting from the highest-level semantic feature, upsampling is realized through parametric transposed convolution. After the output feature is concatenated with the adjacent lower-level feature in the channel dimension, cross-level feature fusion is completed through a convolutional kernel; A cascaded structure is used to process the output of each layer. All-level features are unified to the original input resolution through bilinear interpolation, and are densely concatenated along the channel dimension to form an aggregated feature; A Softmax function is set before the final output layer to normalize the channel dimension and generate a pixel-level probability distribution map; According to the pixel-level probability distribution map, the full connection information of each layer of features is retained to realize the extraction of the spatial distribution information of the lodging winter wheat area.
[0014] Beneficial effects of the present application compared with the prior art: Based on the encoder-decoder structure, the present application constructs a convolutional neural network model. A cross-module feature correction module and an adaptive dual-modal feature fusion module are introduced into the encoder. The cross-module feature correction module corrects the remote sensing image and the DSM elevation image through complementary features between modalities, reduces the loss of small-scale lodging area features in the deep network by dynamically adjusting the receptive field range, thereby improving the segmentation accuracy. The adaptive dual-modal feature fusion module independently extracts the guiding features of the dual-modal of the UAV remote sensing image and the DSM elevation image, and splices and recombines the spectral guiding features and the elevation response features using a cross-grafting strategy to achieve the adaptive mixing of dual-modal features, enhance the discriminative ability of the fused features, and improve the extraction accuracy of the lodging winter wheat area. The multi-scale aggregation decoding module gradually restores the spatial details through learnable deconvolution operations, eliminates the differences in the feature distributions at different scales using channel alignment and group normalization, and adopts a progressive fusion strategy to aggregate the semantic information and texture features at different levels, synchronously enhancing the global context awareness and local boundary sharpening, improving the fuzzy phenomenon of the lodging area boundary extraction, and thus achieving the precise reconstruction of the lodging area from coarse to fine. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic flow chart of a method for extracting the spatial distribution of lodging winter wheat with dual-modal feature fusion provided by an embodiment of the present application; Figure 2 It is a network structure diagram of the DMFCFNet provided by an embodiment of the present application; Figure 3 It is a flow chart of the DMFCFNet training method provided by an embodiment of the present application; Figure 4 The structure diagram of the cross-module feature correction module provided by an embodiment of the present application; Figure 5 The structure diagram of the adaptive dual-modal feature fusion module provided by an embodiment of the present application; Figure 6 The structure diagram of the multi-scale aggregation decoding module provided by an embodiment of the present application; Figure 7 The schematic diagram of the comparison results of different models provided by an embodiment of the present application; Figure 8 The model test result diagram provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The following elaborates on the present solution in combination with the drawings and the specific embodiments.
[0017] Refer to Figure 1 , in the method for extracting the spatial distribution of lodging winter wheat with dual-modal feature fusion in this embodiment, it includes: S101. Establish a dual-stream semantic segmentation network DMFCFNet and train DMFCFNet using the constructed training sample dataset. After testing with the test sample dataset, obtain the optimal DMFCFNet.
[0018] In response to the need for extracting the spatial distribution information of lodged winter wheat, the DMFCFNet network of this application constructs an end-to-end segmentation framework composed of an adaptive dual-modal feature fusion module and a multi-scale aggregation decoder. The overall framework is as Figure 2 shown and can be divided into four parts: dual-modal input, encoder, decoder, and output. Among them, the encoder includes a remote sensing image branch and a DSM elevation image branch. The remote sensing image branch and the DSM elevation image branch include multiple stage residual modules and multiple cross-module feature correction modules. A convolutional layer, a max pooling layer, a batch normalization layer, and a RELU activation function are embedded at the end of each residual module, and each stage residual module is correspondingly connected to a cross-module feature correction module. The output of the cross-module feature correction module is connected to the adaptive dual-modal feature fusion module. The decoder includes a multi-scale aggregation decoding module and a classification module. The multi-scale aggregation decoding module is connected to each adaptive dual-modal feature fusion module in the encoder and is used to receive the fused features output by different adaptive dual-modal feature fusion modules. The multi-scale aggregation decoding module is connected to the classification module.
[0019] In this embodiment, the preprocessed remote sensing image and DSM elevation image are input in parallel into the residual modules of the remote sensing image branch and the DSM elevation image branch for feature extraction, and then feature alignment and complementarity are achieved through the cross-module feature correction module. Then, the corrected remote sensing image features and DSM elevation image features are input into the adaptive dual-modal feature fusion module for adaptive fusion. Finally, the fused feature information at different scales is input into the decoder for segmentation prediction to achieve classification.
[0020] Since Yanzhou District is located in the hinterland of the Huang-Huai-Hai Plain, with a flat and open terrain, the soil is mainly fertile fluvo-aquic soil and cinnamon soil, with deep soil layers, suitable for winter wheat cultivation. The climate belongs to the warm temperate monsoon climate, with distinct seasons, an average annual temperature of 13.5 °C, rain and heat in the same period, and sufficient light and heat resources, providing superior natural conditions for the growth of winter wheat. At the same time, Yanzhou District has rich cultivated land resources and perfect agricultural infrastructure. It is an important grain production base in Shandong Province and has typical regional representativeness. To accurately extract the spatial distribution information of the winter wheat lodging area, Beijiacun in the north of Yanzhou District, where the lodging disaster phenomenon is typical, is selected as the core research area in this embodiment. An unmanned aerial vehicle (UAV) is used to conduct aerial photography on the disaster area of Beijiacun to obtain high-resolution remote sensing images and DSM elevation images. The aerial photography range covers the main farmland distribution area to ensure the integrity of the sample spatial distribution. To avoid cloud occlusion, data collection is carried out on sunny, cloudless days with stable solar radiation intensity. The aerial photography parameters are set as follows: the flight altitude is set to 10 m, the flight speed is 2.6 m / s, the GSD is 0.5 m / pixel, the forward overlap rate is 80%, and the side overlap rate is 70%. The UAV image data includes remote sensing images and DSM elevation images. Among them, the remote sensing images contain three bands: red, green, and blue, and the DSM contains elevation data representing the height information of ground objects.
[0021] After the UAV image data is obtained, geometric correction and radiometric correction are performed on the UAV images to reduce the influence of geometric distortion and illumination differences, ensure that the images correspond to the real ground objects in space, and eliminate the errors caused by changes in flight altitude, angle, and illumination conditions. Subsequently, the image stitching technology is used to seamlessly stitch multiple aerial photography images into a complete winter wheat field map. During the stitching process, a feature matching algorithm is used to automatically identify the overlapping areas to achieve precise alignment. The vector data is marked by the method of manual drawing to generate a marked map containing two categories: lodged winter wheat and non-lodged winter wheat. The preprocessed images and the marked map are segmented using python slicing code to ensure one-to-one correspondence between the original image and the mark, which is convenient for subsequent model training and verification. Based on this, lodged sample data, elevation sample data, and test lodged data are formed. Among them, the lodged sample data and elevation sample data are used to train the model, and the test lodged data is used to test the model. Figure 1 To implement the training of DMFCFNet, the SGD algorithm is used as the training algorithm. The specific training steps include: determining the hyperparameters during the training process
[0022] and and , and initialize each parameter in the network. Use the selected training image - DSM data sample - label pair as the training sample data and input it into DMFCFNet. Use DMFCFNet to perform a forward calculation on the current training data, calculate the loss according to the Loss loss function, use the SGD algorithm to update the parameters of DMFCFNet, complete one training process, repeat the forward calculation on the current training data, and calculate the loss according to the Loss loss function until the loss function is less than the specified expected value or the loss value no longer changes.
[0023] In this embodiment, aiming at the problem of uneven spatial distribution of the lodging winter wheat area in the UAV remote sensing image, a loss function is designed based on the cross - entropy loss and the Dice loss, which is defined as: Among them, is the overall loss, is the cross - entropy loss, is the Dice loss, and are hyperparameters, and the sum of the two is 1.
[0024] The cross - entropy loss calculates the logarithmic difference between the true label and the predicted probability, emphasizing the consistency of the class probabilities. Even in the case of uneven class distribution, it can effectively measure the prediction accuracy of the model for each sample class. The smaller the cross - entropy loss, the higher the matching degree between the predicted probability and the true label. Its calculation formula is.
[0025] Among them, n is the number of samples, is the label value of the i - th sample, is the predicted value of the model for the i - th sample.
[0026] The Dice loss calculates the ratio of the intersection to the union of the predicted region and the actual target region, emphasizing the similarity of the local region. Even in the case of uneven distribution of the target region, it can effectively capture the morphological features of smaller targets. The closer the Dice coefficient is to 1, the better the segmentation effect. Its calculation formula is: Among them, is the label value of the i - th sample, is the predicted value of the model for the i - th sample, and N is the total number of pixel points, which is equal to the number of pixels in a single image multiplied by the batchsize.
[0027] In this embodiment, in order to select appropriate and values, a strategy of gradual modification is adopted, with a smaller the value as the starting point and then gradually increase the value, and each modification takes . Through organizing multiple rounds of training, a set of suitable and values are finally determined.
[0028] As Figure 3 shown, in this embodiment, lodging sample data, elevation sample data, and test lodging data are first prepared. The model is trained using the constructed training sample dataset. During the training process, starting from initializing the model parameters, the samples are input, forward calculations are performed, and the loss is calculated according to the Loss function. The SGD algorithm is used to update the parameters of the model until the loss function is less than the specified expected value or the loss value no longer changes, and the optimal model is obtained. After the model is determined, field extraction is performed, then spatial distribution extraction is performed for winter wheat, followed by accuracy evaluation. The performance of the model is analyzed through the results of the accuracy evaluation, the weight file is saved. After the training is completed, the model is tested using the test lodging data, and the spatial distribution information of the lodging winter wheat is output.
[0029] S102, Parallelly input the preprocessed remote sensing image and the DSM elevation image into the DMFCFNet to obtain the semantic texture feature information of the remote sensing image at different scales and the spatial structure feature information of the DSM elevation image.
[0030] After parallelly inputting the preprocessed remote sensing image and the DSM elevation image into the DMFCFNet, the input layer of the remote sensing image branch uses a 7×7 large kernel convolution to construct a wide receptive field, effectively capturing the low-frequency texture information in the image, identifying the features of the lodging winter wheat, and reducing the image size through downsampling to reduce the subsequent calculation amount. The first-stage residual module Res1 retains the high-resolution details through skip connections and compresses the feature map with a 3×3 max pooling at the end, retaining the key information and effectively suppressing the noise. The second to fourth-stage residual modules Res2–Res4 use 1×1–3×3–1×1 compression–extraction–expansion residual units to gradually mine the deep semantic features, and dilated convolutions with dilation rates of 2 and 4 are respectively applied in Enc3 and Enc4 to expand the receptive field while maintaining the resolution of the feature map, thereby enhancing the cross-region context modeling ability for scattered lodging patches.
[0031] The DSM elevation image branch adjusts the number of convolutional kernel channels in the input layer to 1 for single-channel elevation data, and reveals the spatial correlation between terrain undulation and crop lodging through elevation gradient features. The remaining structures of this branch are the same as those of the remote sensing image branch, including the same downsampling strategy and compression–extraction–expansion residual units. In addition, a batch normalization layer BN and a RELU activation function are embedded at the end of each residual block in both branches, which can eliminate the distribution differences between channels and enhance the nonlinear expression ability.
[0032] S103. Feature correction and fusion are performed on the semantic texture feature information of the remote sensing image and the spatial structure feature information of the DSM elevation image to obtain the fused feature information.
[0033] In this embodiment, to solve the problem of non-linear association of the lodging area in the multi-modal feature space, a cross-modal feature correction module CFCM is introduced after each level of feature extraction in the model. This module constructs a unit with CPCA attention as the core for the remote sensing image feature map and the DSM elevation image feature map at the same level, integrates a learnable weight assignment strategy, dynamically models the complementary association between modalities while strengthening the discriminative ability of single-modal features. The structure of the CFCM module is as Figure 4 shown, including: a channel attention module, a standard convolution module, a depthwise separable convolution layer, a spatial attention module, and a cross-modal cross-attention module.
[0034] First, global average pooling and max pooling are respectively performed on the input remote sensing image feature map and the DSM elevation image feature map to capture the mean information and extreme value responses at the channel level. Then, two layers of 1×1 convolutions are performed for non-linear transformation to generate channel attention weights, and the semantic clues of the two poolings are fused by element-wise addition. Finally, after passing through the sigmoid activation function, adaptive weighting in the channel dimension is realized, and the corrected features and .
[0035] The corrected features and Processing is carried out. The 5×5 convolution is used to capture the basic spatial relationships in the local neighborhood. The asymmetric convolution kernel reduces the computational amount by decomposing the spatial dimension while covering longer spatial dependencies, effectively extracting structural features with different aspect ratios and adapting to the morphological diversity of lodged winter wheat. All spatial convolution layers adopt grouped convolution to enhance local feature interaction. Finally, multi-scale features are fused through residual connections, and a spatial attention map is generated through 1×1 convolution and the RELU activation function to achieve the spatial structure optimization of the input features.
[0036] To achieve co-optimization between modalities, the corrected features and are dynamically complemented through cross-modal cross-attention CrossAttention. First, the bimodal features of the remote sensing image feature map and the DSM elevation image feature map are concatenated, and cross-modal associations are extracted through 3×3 convolution and the RELU activation function. Finally, spatial weights are generated by 1×1 convolution and , which are used to identify the dependence degree of each pixel point on the features of the other modality. Finally, two learnable parameters and are used to perform weighted fusion on the bimodal features to generate output features. The calculation formula is: where, , respectively represent the remote sensing image and the DSM elevation image, and represent the learnable parameters, , represent the corrected remote sensing image features and the corrected DSM elevation image features, and represent the spatial weights, represents element-wise multiplication.
[0037] In addition, to further reduce the offset problem caused by the possible differences between modalities during the training process, a Dropout layer is added after the output of each layer of the feature map, which alleviates the overfitting problem and improves the robustness and generalization ability of the model.
[0038] The Adaptive Bimodal Feature Fusion Module AMFM realizes the adaptive fusion of remote sensing image and DSM elevation image features through a structured cross-modal interaction and dynamic spatial fusion mechanism. The structure diagram is as shown in Figure 5As shown in the figure. The AMFM module adopts a dual-branch parallel architecture to process the corrected remote sensing image features and DSM elevation image features respectively. Each branch uses an independent 1×1 convolution to expand the input channel to 3 times the hidden channel, and then divides it into three groups of feature maps along the channel dimension, namely the guide feature (locating the key area of cross-modal information interaction), the association feature (establishing the cross-modal feature similarity measurement) and the content feature (carrying the modality-specific information to be transmitted). Then, two-way interaction is achieved through cross-modal feature splicing. The guide feature of the remote sensing image is spliced with the association feature and content feature of the DSM along the channel dimension to form a fusion feature, so that the spectral feature integrates the terrain structure information. At the same time, the DSM guide feature is reversely spliced with the association feature and content feature of the remote sensing image, and its texture detail features are introduced. Through cross-modal feature exchange, the complementary association between the two modalities is modeled. The fused high-dimensional features are compressed to the hidden channel dimension by 1×1 convolution respectively, and a compact representation is generated to reduce the computational complexity. To avoid the loss of small-scale lodging features during dimensionality reduction, the model maps the compact representation back to the original number of channels through a projection layer and adds it pixel by pixel to the original input features.
[0039] In order to adapt to the scattered patches and large areas of winter wheat that are mixed in the lodging area, a spatial dynamic weight map is finally generated through a spatial adaptive fusion network to achieve pixel-level adaptive adjustment of the modal fusion ratio. The network concatenates the features after residual enhancement along the channel, and successively passes through 3×3 depth-separable convolution, RELU activation function and 1×1 convolution to output a weight tensor. The Softmax function is applied to the tensor along the channel dimension to obtain a spatial weight map, and the final fusion feature is calculated by weighted summation. This mechanism can dynamically adjust the fusion ratio according to the local context, enhance the weight of remote sensing images in areas with depth information defects, and rely on depth features in areas with weak textures.
[0040] S104, performing segmentation and prediction on the fused feature information of multiple scales to extract the spatial distribution information of the winter wheat lodging area.
[0041] In this embodiment, the decoder part adopts a progressive feature pyramid architecture, which focuses on solving the problem of blurred boundaries and missed detection of small targets caused by scale differences in the lodging area through multi-stage feature interaction and spatial information reconstruction. Figure 6 The figure shows the structure of the multi-scale aggregation decoding module MAE. The decoder input is the multi-scale feature map of four different levels extracted by the backbone network after the adaptive bimodal feature fusion module, covering from low-level detail features to high-level semantic features.
[0042] First, the channel dimension alignment layer AlignConv is used to align the multi-level input features in the channel dimension. A parallel 1×1 convolutional layer is adopted to uniformly map the features of each layer. Group normalization and the GELU activation function are introduced to eliminate the cross-modal feature distribution differences and reduce the computational complexity. In the feature fusion part, MAE adopts a top-down information fusion TopDown path. Starting from the highest-level semantic features, 2x upsampling is achieved through parametric transposed convolution. After the output features are concatenated with the adjacent lower-level features in the channel dimension, cross-level feature fusion is completed through a 3×3 convolutional kernel. This process gradually injects high-level semantic information into the low-level features through a stacked structure to form a feature pyramid with rich context information. Among them, the transposed convolution uses a 3×3 kernel size and a stride of 2, and the padding parameters are used to ensure the accurate expansion of the feature map size. Compared with the traditional bilinear interpolation method, the learnable parameter mechanism can adaptively correct the feature distortion in the upsampling process. The multi-scale feature aggregation module uses a cascaded structure to process the outputs of each layer of the pyramid. All-level features are unified to the original input resolution through bilinear interpolation, and densely concatenated along the channel dimension to form aggregated features. After being compressed by a 3×3 convolution, these features are directly mapped to the target category space through a 1×1 convolution. In addition, a Softmax function is set before the final output layer to normalize the channel dimension and generate a pixel-level probability distribution map. By retaining the full connection information of each layer of features, the model's representation ability for multi-scale targets is enhanced, effectively maintaining the edges when extracting the areas of lodged winter wheat.
[0043] To verify the above method, the following experiments were designed for verification.
[0044] The U-Net, DeepLabv3+, RedNet, and MMFNet models were selected as comparison models to conduct comparative experiments to analyze the influence of different network structures on the accuracy of generating the DSM of the winter wheat planting area. The main features of each model are shown in Table 1.
[0045] Four metrics, namely Accuracy, Precision, Recall, and F1-score, were selected to evaluate the accuracy of the experimental results.
[0046] In the comparative experiment of this embodiment, a computer equipped with an NVIDIA GeForce RTX 4060 GPU was used as the test platform. The deep learning framework was selected as PyTorch. The source codes of the comparison models were derived from the open literature, and the rest of the codes were written in the Python language. In addition, with the help of open-source libraries such as GDAL and OpenCV, the image processing programs required for the comparative experiment were developed. To verify the effectiveness of different modules, four groups of ablation experiments were set up in this embodiment, as shown in Table 2. In Experiment 1, the feature maps of remote sensing images and DSM elevation image data were directly spliced in the channel dimension, and a decoder architecture containing only an upsampling layer and a convolutional layer was adopted. In Experiment 2, CFCM was used in the encoder part. In Experiment 3, AMFM was used to replace the simple splicing and fusion of remote sensing images and DSM elevation image information. In Experiment 4, a multi-scale aggregation decoder was used on the basis of the existing architecture. This experimental design can effectively verify the influence of each module on the extraction accuracy of the spatial distribution of lodged winter wheat. The ablation experiment data shows that adding only the CFCM module (Experiment 2) can make the accuracy, precision, recall rate, and F1 score reach 92.3%, 88.5%, 89.2%, and 88.6% respectively, verifying the improvement effect of feature correction on the problems of blurred boundaries and missed classification of small targets. Experiment 3 is lower than Experiment 4 using the multi-scale aggregation decoder in terms of evaluation indicators, indicating that there are obvious limitations in the single-scale reconstruction mechanism. Especially in small-scale lodged areas, due to the gradual attenuation of the feature resolution during bilinear interpolation upsampling, fine-grained textures cannot be retained, resulting in misclassification of the lodged winter wheat area. On the basis of Experiment 3, after introducing the MAE module, the extraction precision rate of the model in the lodged patch area reaches 89.7%. Its advantage stems from the collaborative optimization of high-level semantic features and low-level detail features. High-level features provide the global distribution prior of the lodged area, while low-level features reconstruct local geometric patterns through deconvolution. In addition, the progressive fusion strategy enhances the model's perception ability of the heterogeneous distribution in the lodged area by densely splicing multi-scale features.
[0047] After the extraction is completed on the test data using all comparison methods, 5 image patches are randomly selected from the test results of Lijia Village for visually demonstrating the extraction effects of the DMFCFNet method and the comparison methods. Figure 7 The results generated by the five models in five regions are presented. Among them, Figure 7 (a) in it is the original remote sensing image; Figure 7 (b) in it is the labeled image, Figure 7 (c) in it is the recognition result of the DMFCFNet model, Figure 7 (d) in it is the recognition result of the MMFNet model, Figure 7 (e) in it is the recognition result of the REDNet model, Figure 7 (f) in it is the recognition result of the DeepLabv3+ model, Figure 7In (g), the recognition result of the U-Net model is shown. The green and black areas represent the extracted lodging areas and non-lodging areas respectively. The results show that the extraction result of the U-Net model is the worst, with blurred edges and fractures in the lodging areas, and widespread misrecognition and incorrect recognition. The DeepLabv3+ model improves the regional integrity through multi-scale feature fusion, and the misrecognition rate is reduced, but the boundary segmentation of the transition area is still inaccurate, and obvious missing detections of small fragmented patches occur. Although the RedNet model fuses UAV images and DSM elevation data, the bimodal features fail to cooperate effectively, resulting in increased roughness of edge segmentation and being vulnerable to shadow interference, with limited ability to extract small target lodging areas. The MMFNet model introduces a 3D convolutional block attention module based on RedNet, improving the integrity and boundary clarity of the extracted lodging areas, but still lacking in the performance for small target lodging winter wheat areas. In contrast, the DMFCFNet model proposed in this embodiment better improves the integrity and boundary clarity of lodging area extraction, effectively suppressing misrecognition in complex scenarios, and achieving accurate detection even for small fragmented patches, with the best comprehensive performance.
[0048] Table 3 presents the evaluation metrics of all comparison methods on the entire test set. It can be seen that in the test results of the Lijiazhuang Village dataset in Yanzhou District, the metrics of DMFCFNet are significantly better than those of other comparison methods.
[0049] As can be seen from Table 3, the F1-score metric of the DMFCFNet model is improved by approximately 5.3, 3.2, 2.3, and 1.2 percentage points compared to the U-Net, Deeplabv3+, RedNet, and MMFNet models respectively. Considering all metrics, the DMFCFNet model can be well applied to the task of extracting the spatial distribution of lodging winter wheat in UAV remote sensing images.
[0050] Based on the constructed DMFCFNet model, in this embodiment, 0.5-meter high-resolution UAV remote sensing image data is used, and the data preparation of UAV data and DSM data for the eastern experimental area of Wangfuzhuang Village is completed with the same data preprocessing process as the Beilijiazhuang Village dataset. Based on the model weights obtained from the training of Lijiazhuang Village, a validation experiment is carried out on the lodging areas of winter wheat in Wangfuzhuang Village to systematically evaluate the generalization performance and extraction accuracy of the model on different datasets. Figure 8 For the extraction result of lodging winter wheat in Wangfuzhuang Village, the DMFCFNet method effectively distinguishes the categories of non-lodging winter wheat, construction land, etc. in the eastern part of Wangfuzhuang Village from the lodging winter wheat areas, and effectively removes the influence of infrastructure such as field roads and trees, accurately completing the task of extracting lodging winter wheat, verifying the effectiveness and reliability of the DMFCFNet model in different scenarios.
[0051] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent the cases of A existing alone, A and B existing simultaneously, and B existing alone. Where A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c may be single or multiple.
[0052] As described above, the foregoing is only a specific embodiment of the present application. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. The protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for extracting the spatial distribution of lodging winter wheat by dual-modal feature fusion, characterized in that, Including: Establish a dual-stream semantic segmentation network DMFCFNet and train DMFCFNet using the constructed training sample dataset, and obtain the optimal DMFCFNet after testing with the test sample dataset; Parallelly input the preprocessed remote sensing image and DSM elevation image into DMFCFNet to obtain the semantic texture feature information of the remote sensing image at different scales and the spatial structure feature information of the DSM elevation image; Perform feature correction and fusion on the semantic texture feature information of the remote sensing image and the spatial structure feature information of the DSM elevation image to obtain the fused feature information; Perform segmentation prediction on the fused feature information at multiple different scales to achieve the extraction of the spatial distribution information of the winter wheat lodging area.
2. The method for extracting the spatial distribution of lodging winter wheat with dual-modal feature fusion according to claim 1, wherein The DMFCFNet includes: an encoder and a decoder. The encoder includes a remote sensing image branch and a DSM elevation image branch. The remote sensing image branch and the DSM elevation image branch include multiple stage residual modules and multiple cross-module feature correction modules. Each stage residual module is correspondingly connected to a cross-module feature correction module, and the output of the cross-module feature correction module is connected to the adaptive dual-modal feature fusion module; The remote sensing image and the DSM elevation image are parallelly input into the residual modules of the remote sensing image branch and the DSM elevation image branch, and feature alignment and complementarity are achieved through the cross-module feature correction module; The adaptive dual-modal feature fusion module adaptively fuses the corrected remote sensing image features and DSM elevation image features; The decoder includes a multi-scale aggregation decoding module and a classification module. The multi-scale aggregation decoding module is connected to each adaptive dual-modal feature fusion module in the encoder for receiving the fused features output by different adaptive dual-modal feature fusion modules. The multi-scale aggregation decoding module is connected to the classification module. The decoder is used to perform segmentation prediction on the fused feature information at different scales to achieve classification.
3. The method for extracting the spatial distribution of lodged winter wheat with bimodal feature fusion according to claim 1, wherein The establishment of the dual-stream semantic segmentation network DMFCFNet and the training of DMFCFNet using the constructed training sample dataset include: Determine the hyperparameters during the training process and initialize each parameter in the network; Input the selected training image-image-DSM data sample-label pair as the training sample data into DMFCFNet; Perform a forward calculation on the current training data using DMFCFNet and calculate the loss according to the Loss loss function; Update the parameters of DMFCFNet using the SGD algorithm until the loss function is less than the specified expected value or the loss value no longer changes.
4. The method for extracting the spatial distribution of lodged winter wheat with bimodal feature fusion according to claim 3, wherein, The calculation formula of the Loss loss function is: Among them, is the cross-entropy loss, is the Dice loss, and are hyperparameters, n is the number of samples, is the label value of the i-th sample, is the predicted value of the model for the i-th sample, and N is the total number of pixel points.
5. The method for extracting the spatial distribution of lodged winter wheat with bimodal feature fusion according to claim 1, wherein, The parallel input of the preprocessed remote sensing image and DSM elevation image into DMFCFNet to obtain the semantic texture feature information of the remote sensing image at different scales and the spatial structure feature information of the DSM elevation image includes: After the pre - processed remote - sensing image and the DSM elevation image are input into the DMFCFNet in parallel, the input layer of the remote - sensing image branch constructs a wide receptive field using large - kernel convolution to capture the low - frequency texture features in the input remote - sensing image, and downsamples to reduce the size of the remote - sensing image; For the single - channel elevation data of the DSM elevation image branch, the number of channels of the convolution kernel in the input layer is adjusted, the spatial correlation between terrain undulation and crop lodging is revealed through elevation gradient features, and the size of the DSM elevation image is reduced by downsampling; The first - stage residual module retains the high - resolution details of the remote - sensing image and the DSM elevation image through skip connections, compresses the feature map with max - pooling at the end, and extracts deep semantic features through the extended residual unit.
6. The method for extracting the spatial distribution of lodging winter wheat by bimodal feature fusion according to claim 1, characterized in that The semantic texture feature information of the remote - sensing image and the spatial structure feature information of the DSM elevation image are feature - corrected and fused to obtain the fused feature information, including: Global average pooling and max - pooling are respectively performed on the input remote - sensing image and DSM elevation image. After capturing the channel - level mean information and extreme - value responses, non - linear transformation is performed through a convolution module to generate channel attention weights; Fuse the semantic clues of the two pooling methods by element-wise addition, and finally achieve adaptive weighting in the channel dimension through the sigmoid activation function, and output the corrected features and ; The corrected features are processed using multi-scale depthwise separable convolutions with standard convolutions and asymmetric convolution kernels and processed; The corrected features and achieve dynamic complementarity through the cross-modal cross-attention module; The adaptive dual - modal feature fusion module realizes the adaptive fusion of the features of the remote - sensing image and the DSM elevation image through a structured cross - modal interaction and dynamic spatial fusion mechanism.
7. The method for extracting the spatial distribution of lodged winter wheat by bimodal feature fusion according to claim 6, wherein The corrected features and achieve dynamic complementarity through a cross-modal cross-attention module, including: The dual - modal features of the remote - sensing image and the DSM elevation image are concatenated; Extract cross-modal associations through 3×3 convolution and the ReLU activation function, and generate spatial weights by 1×1 convolution and ; Using two learnable parameters and perform weighted fusion on the bimodal features to generate output features, and the calculation formula is as follows: Among them, and represent a remote sensing image and a DSM elevation image respectively, and represent learnable parameters, and represent the features of the corrected remote sensing image and the features of the corrected DSM elevation image, and represent spatial weights, represents element-wise multiplication.
8. The method for extracting the spatial distribution of lodging winter wheat by bimodal feature fusion according to claim 6, characterized in that, The adaptive dual - modal feature fusion module realizes the adaptive fusion of the features of the remote - sensing image and the DSM elevation image through a structured cross - modal interaction and dynamic spatial fusion mechanism, including: The adaptive dual - modal feature fusion module adopts a dual - branch parallel architecture to process the corrected remote - sensing image features and DSM elevation image features respectively; Each branch uses an independent convolution module to expand the dimension of the input channels and divides them into three groups of feature maps: guiding features, associated features, and content features along the channel dimension; The guiding features of the remote - sensing image are concatenated with the associated features of the DSM elevation image and the content features of the DSM elevation image along the channel dimension to form the first fusion feature; At the same time, the guiding features of the DSM elevation image are concatenated with the associated features of the remote - sensing image and the content features of the remote - sensing image in reverse to form the second fusion feature; The first fusion feature and the second fusion feature are respectively compressed to the hidden channel dimension through convolution to generate the first compact representation and the second compact representation; The first compact representation and the second compact representation are mapped back to the original number of channels through a projection layer and added to the input corrected remote - sensing image features and DSM elevation image features pixel - by - pixel; A spatial dynamic weight map is generated through a spatial adaptive feature fusion network to realize the adaptive fusion of the features of the remote - sensing image and the DSM elevation image.
9. The method for extracting the spatial distribution of lodged winter wheat with dual-modal feature fusion according to claim 1, characterized in that, The extraction of the spatial distribution information of the winter wheat lodging area is realized by segmenting and predicting the fused feature information of multiple different scales, including: The fused feature information of multiple different scales is aligned in the channel dimension, and each layer of features is uniformly mapped using a parallel convolutional layer; After eliminating the cross-modal feature distribution differences through group normalization and the GELU activation function, a top-down information fusion path is adopted. Starting from the highest-level semantic features, upsampling is achieved through parametric transposed convolution. After the output features are concatenated with adjacent lower-level features along the channel dimension, cross-level feature fusion is completed through convolutional kernels; A cascaded structure is used to process the outputs of each layer. All-level features are unified to the original input resolution through bilinear interpolation, and dense concatenation is performed along the channel dimension to form aggregated features; Before the final output layer, a Softmax function is set to perform normalization processing on the channel dimension to generate a pixel-level probability distribution map; According to the pixel-level probability distribution map, the full connection information of each layer of features is retained to achieve the extraction of the spatial distribution information of the area of lodged winter wheat.
Citation Information
Patent Citations
High-resolution remote sensing image land coverage classification method and device and storage medium
CN117036936A
Bimodal target detection method and system
CN117809202A
Unmanned aerial vehicle wheat lodging image segmentation method based on DeeplabV3 + and Unet network architecture
CN118172688A
Ultrahigh-resolution remote sensing image segmentation method based on frequency domain information fusion
CN119559200A
Edge-guided winter wheat fine space distribution remote sensing extraction method
CN119762986A
Cited By
Pipeline backfilling settlement monitoring method and system based on image recognition
CN120747863A
Cultivated land change detection method based on cross-modal remote sensing data
CN121330499A
Method for detecting cultivated land change based on cross-modal remote sensing data
CN121330499B