Wheat lodging region recognition model establishment method and application

Through the UssNet model combining GeLu activation function and Focal loss function, the problems of inefficiency and insufficient accuracy of traditional wheat lodging area recognition methods are solved, and efficient and accurate wheat lodging area recognition is achieved.

CN120388264APending Publication Date: 2025-07-29HENAN AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510320678.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

传统的小麦倒伏区域识别方法处理效率低下、成本高且分割识别精度不高,尤其在轻微倒伏时精度不足。

Method used

The UssNet model is adopted, combined with the MCC module, CCR module, SSModule module and RestBlock module, and the GeLu activation function and Focal loss function are used to identify and segment wheat lodging areas. The spatial similarity perception ability of the model is improved through the state space perception module, and the sample distribution is optimized through data augmentation and model training.

Benefits of technology

提高了小麦倒伏区域的识别精度和效率,实现了更高的分割精度和更好的性能平衡,能够有效识别轻微倒伏区域。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388264A_ABST
    Figure CN120388264A_ABST
Patent Text Reader

Abstract

The invention discloses a wheat lodging region recognition model establishment method and application, and mainly solves the technical problems of low processing efficiency, high cost and low segmentation recognition precision of a traditional recognition method. Through the UssNet model combining context and space perception, the advantages of semantic segmentation of UNet and global context perception of State Space Models (SSM) are brought into full play, and the method can be well applied to a wheat lodging semantic segmentation task. And a local auxiliary mechanism is combined with SSM, and feature mapping is selectively scanned from different directions, so that a long sequence can be linearly calculated, and global context image information can be extracted. In addition, a Focal loss function is introduced to solve the problems of wheat lodging category imbalance and lodging small region identification. In order to solve the dead ReLu phenomenon, GeLu is adopted as an activation function, and the over-fitting suppression capability of ReLu is reserved, so that the optimal recognition effect can be achieved. By paying attention to image global information, the UssNet model can effectively improve the recognition precision of the wheat lodging region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of agricultural disaster assessment, and particularly relates to a method for establishing a wheat lodging area recognition model and its application. Background Art

[0002] Crop lodging refers to the tilting of the stem caused by disasters and is regarded as one of the main negative factors restricting crop yields. Lodging is usually caused by natural disaster conditions such as strong winds and short-term heavy precipitation, as well as inappropriate crop management measures (such as high planting density, excessive application of nitrogen fertilizer during cultivation, etc.). Moreover, lodging can occur at different growth stages of wheat, especially lodging in the middle and late stages has a more adverse impact on wheat yields. Therefore, quickly and accurately obtaining the wheat lodging area is crucial for estimating yield losses, post-disaster management measures, and supporting breeding decisions.

[0003] Currently, the monitoring methods for wheat lodging areas mainly include ground manual surveys and remote sensing monitoring. The manual survey method involves researchers going deep into the field for visual inspections and evaluating the observed lodging situations. This method is time-consuming, costly, inefficient, and can cause secondary damage to the wheat. Remote sensing monitoring, on the other hand, provides a fast and non-destructive alternative method for estimating lodging over a larger area. With the explosive increase in the demand for precise agricultural management, drones are increasingly used in high-resolution scenarios, and various studies on predicting lodging using RGB, radar (LIDAR), and multispectral sensors installed on drone systems have provided valuable insights into the challenges faced by other crops. For example, a method for monitoring wheat lodging areas based on a hybrid watershed algorithm and an adaptive threshold segmentation algorithm is known to the inventor. This method is based on drone remote sensing images and monitors wheat lodging during the filling period, with a lodging extraction accuracy of 93.58%. In addition, the inventor is also aware of a lodging detection model constructed based on the spectral reflectance, vegetation indices, texture, and color features of drone images. The research results show that the red edge and near-infrared bands can effectively distinguish lodging areas, with an overall accuracy of over 90%.

[0004] However, in the process of implementing the technical solutions in the embodiments of the present application, the inventors found that although the above traditional feature extraction methods can effectively extract crop lodging information using remote sensing images, they rely on the selection and quantity of samples, have a relatively high cost in the process of processing data, and are inefficient in processing long sequences of remote sensing images and have a low accuracy in segmenting lodging areas.

[0005] The information disclosed in this background art section is only for deepening the understanding of the background art of the present disclosure and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art. Summary of the Invention

[0006] In view of at least one of the above technical problems, the present disclosure provides a method for establishing a wheat lodging area recognition model and its application, mainly solving the technical problems of low processing efficiency, high cost, and low segmentation recognition accuracy of traditional recognition methods.

[0007] According to one aspect of the present disclosure, a method for establishing a wheat lodging area recognition model is provided, which includes the following steps: (1) Obtain wheat lodging images by a multispectral drone under clear and cloudless weather conditions; (2) Preprocess the wheat lodging images obtained by the multispectral drone to obtain a sample data set; and correspondingly divide the data set into a training set, a validation set, and a test set; (3) Establish a UssNet model including an MCC module, a CCR module, an SSModule module, and a RestBlock module; the encoder of the UssNet model is correspondingly cross-composed of the MCC module and the CCR module, and the decoder of the UssNet model corresponds to the structure of the encoder; (4) Use the Focal loss function and train the UssNet model based on the sample data set.

[0008] In some embodiments of the present disclosure, in step (2), before the preprocessing, the wheat lodging images with tilted fields of view and irrelevant shadows are first excluded; the preprocessing includes image cropping for cropping the wheat lodging images through a rectangular window of a specific size to generate the images of the sample data set and image annotation for annotating the contours of the wheat lodging areas.

[0009] In some embodiments of the present disclosure, in step (2), the images belonging to the wheat lodging category in the sample data set are subjected to angle transformation and data augmentation processing, and all the processed extended data are correspondingly divided into the training set.

[0010] In some embodiments of the present disclosure, in step (3), the MCC module performs convolution processing on the model input and then uses the GeLu activation function to nonlinearly rectify the features; and the MCC module further includes a normalized BatchNorm2D layer.

[0011] In some embodiments of the present disclosure, in step (3), the SSModule module takes the features nonlinearly rectified by the GeLu activation function as input, uses a dual-path processing data including a linear embedding layer respectively as output, and sequentially performs depth convolution, SiLu activation, SSM module, and layer normalization processing along one of the path branches.

[0012] In some embodiments of the present disclosure, in the step (3), the CCR module includes two convolutional layers and uses the GeLu activation function.

[0013] In some embodiments of the present disclosure, in the step (3), the RestBlock module is respectively introduced into the MCC module and the CCR module.

[0014] In some embodiments of the present disclosure, in the step (4), the model training adopts the dropout random discard algorithm, the hyperparameter is set to 0.1, the training epoch is 300 rounds, the batch is 4, the initial learning rate is 0.001, the learning rate update strategy is the cosine annealing strategy, the weight decay coefficient is 0.001, and the optimization algorithm is AdamW.

[0015] According to another aspect of the present disclosure, there is provided a method for identifying wheat lodging areas, which identifies wheat lodging areas based on the UssNet model trained correspondingly according to the above method.

[0016] One or more technical solutions provided in the embodiments of the present application have at least any one of the following technical effects or advantages: 1. By adding a state space perception module, the UssNet model has the advantage of performance improvement compared with traditional CNN models. In addition, the state space perception model has the ability to perceive spatial similarity that traditional CNN models do not have, breaking through the limitation that the receptive field of traditional CNN models is limited by the size of the convolutional kernel, so that the UssNet model can perceive the elevation intervals where different targets are located.

[0017] 2. The UssNet model uses the GeLu activation function to solve the problem of "neuron death" caused by the traditional ReLu activation function during training. Although as long as there are negative numbers, the neuron will become 0 in subsequent calculations, but some negative neurons can play a role in subsequent calculations; by using GeLu as the activation function, it can slightly allow the appearance of negative values and retain the overfitting suppression ability of ReLu, improving the model performance.

[0018] 3. The UssNet model introduces the Mamba state space SSM module with context awareness. The Mamba model effectively identifies the correlation between pixels through two MLP modules. On this basis, a 2D convolutional module is added to extract features from the pixel correlation. This approach improves the quality of the feature layer compared to simple convolution operations. Under these conditions, the segmentation accuracy of UssNet is higher because the UssNet model can learn the global features of the image more effectively. In addition, after adding the SSMoudle module to the UssNet model, with the entire image as the receptive field, it directly affects the model's global perception ability and the understanding ability of the lodging area, thereby improving the model's accuracy and significantly enhancing the performance of the network.

[0019] 4. The UssNet model can make full use of multi-scale feature fusion and context information capture technology to accurately locate the detailed features of the lodging edge, demonstrating its superiority in feature expression and spatial information utilization, and achieving a good balance between performance and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is an image of the lodging part of wheat obtained by a multi-spectral unmanned aerial vehicle in an embodiment of the present disclosure.

[0021] Figure 2 It is a schematic diagram of the network structure principle of the UssNet model in an embodiment of the present disclosure.

[0022] Figure 3 It is a comparison of the segmentation results of each model in an embodiment of the present disclosure.

[0023] Figure 4 It is a comparison of the performance of each model on the complete remote sensing image of wheat lodging in an embodiment of the present disclosure; among them, (a) is the original image, (b) is UssNet, (c) is UNet, (d) is SegNet, (e) is DeepLabv3+, and (f) is ViT.

[0024] Figure 5 It is a comparison of the training results of each model during the model training process in an embodiment of the present disclosure; among them, (a) is the mIoU of wheat lodging, and (b) is the pixel accuracy.

[0025] Figure 6 It is the effect of UssNet in the aspect of regional boundary recognition in an embodiment of the present disclosure; where a to e are different lodging areas respectively.

[0026] Figure 7 It is the training result of the cross-entropy loss function in an embodiment of the present disclosure.

[0027] Figure 8This is the training result of the Focal loss function in an embodiment of the present disclosure.

[0028] Figure 9 This is the result of some ablation experiments in an embodiment of the present disclosure. Detailed implementation manners

[0029] To better understand the technical solution of the present application, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0030] To solve the technical problems that the traditional method for identifying wheat lodging areas is highly dependent on the selection and quantity of samples, has a high cost in the process of processing data, and has low accuracy in identifying slight lodging, this example discloses a method for establishing a wheat lodging area identification model, which specifically includes the following steps: (1) Obtain wheat lodging images by a multispectral drone under clear and cloudless weather conditions.

[0031] In late May, a certain experimental field in Yuanyang County, Xinxiang City, Henan Province and Huiji District, Zhengzhou City encountered abnormal weather such as strong winds and heavy rains during the wheat filling stage and maturity stage, resulting in wheat lodging. Therefore, in this embodiment, the above experimental field with wheat lodging is used as the data source to identify the wheat lodging area in the field; among them, the lodging wheat varieties include Xinmai 26, Yunong 908, Yunong 186, Yangmai 15, and Kexing 3302. After the lodging occurred, data collection was carried out respectively under clear and cloudless weather, and a DJI Mavic3M multispectral drone was used to obtain the wheat lodging images.

[0032] Specifically, set the flight height of the multispectral drone from the ground to be 12m, the flight speed to be 1.5m / s, and the bidirectional overlap rate to be 80%. Refer to Figure 1 , Figures (a) and (b) in the figure are images of the experimental field in Yuanyang County, and Figure (c) in the figure is an image of the experimental field in Huiji District. It can be clearly seen from some of the wheat lodging images obtained by the multispectral drone that compared with the wheat growing upright, the lodging area shows obvious changes in texture and color, and the texture of the lodging area is more chaotic and appears particularly prominent visually. In addition, the contrast between light and dark in the lodging area is also more obvious than that in the non-lodging area.

[0033] (2) Preprocess the wheat lodging images obtained by the multispectral drone to obtain a sample data set; and divide the data set into a training set, a validation set, and a test set correspondingly.

[0034] In this embodiment, when the multispectral drone collects data images, it takes pictures automatically at equal intervals. Therefore, there are many unusable images in the acquired images. Since the size of a single image in this example is 5472×3648 pixels, which is relatively large, to improve the data processing efficiency and avoid excessive occupation of computer memory and RAM by subsequent model operations, in this example, the data images in the wheat lodging dataset are preprocessed first, which includes image cropping and image annotation.

[0035] Specifically, first, the original images with tilted fields of view and irrelevant shadows are removed from the data images obtained by the multispectral drone, and then each original image is cropped using a rectangular window of 512×512 pixels to generate samples. In this example, a total of 1326 sample images are obtained after cropping as the sample dataset. Then, the wheat images in the sample dataset are divided into two categories: lodged and non-lodged according to whether they are lodged. In this example, polygon annotation in the Labelme annotation software is used to perform pixel-level annotation on the wheat lodging area, and when annotating, the bounding box of the annotation is made as close as possible to the contour of the lodging area to maximize the annotation accuracy.

[0036] To obtain the optimal sample size, in this embodiment, enhanced data processing is performed on the basis of the 1326 RGB images obtained in the early stage. At the same time, considering that the distribution of the number of pixels in lodged and non-lodged images is uneven, in this case, the model may tend to predict more categories and thus ignore fewer categories. To balance the pixels of the two types of samples, in this embodiment, the images containing the lodged category are subjected to angular transformations including 45° rotation, horizontal flipping, and vertical flipping, and the Mosaic data augmentation method is used for augmentation. The Mosaic data augmentation used in this example selects multiple images from the dataset and then performs basic data augmentations such as scaling and gamut change on the multiple images respectively. The dataset expanded by angular transformation and data augmentation processing contains 5052 images of 512×512 pixels, among which the proportions of lodged and non-lodged pixels are 49.66% and 51.34% respectively. Finally, the dataset is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1 (i.e., 3500 images, 1000 images, and 500 images), where the validation set and the test set do not include the data of the enhanced images.

[0037] (3) Establish a UssNet model including an MCC block, a CCR block, an SSModule block, and a RestBlock block.

[0038] In this embodiment, refer to Figure 2, the UssNet model has a U-shaped network structure including an encoder and a decoder, and the encoder and decoder are symmetric. Specifically, in this example, the UssNet model integrates the MCC (Mamba Convolution Convolution) module, the CCR (Convolution Convolution Resnet) module, the SSModule (State Space Module) module, and the RestBlock module. Among them, the encoder part of the UssNet model is composed of the cross of the MCC module and the CCR module. The MCC module is used for feature extraction of sample correlation, and the CCR module is used for the extraction of image texture and graphics. In this embodiment, the encoding layer contains 2 MCC modules and 2 CCR modules, a total of 4 encoding modules. Each time a feature extraction is performed, a max pooling layer is added as the downsampling of the model, and the size of the downsampled image is reduced by (1 / 2) n , where n is the number of layers. In addition, the structures of the decoder and the encoder correspond to each other, and the upsampling is used instead of the downsampling part of the encoder for the decoder. In addition, in order to improve the MIou accuracy of the Ussnet model and the fineness of the segmentation edge, in this embodiment, bilinear interpolation is used for upsampling, thereby obtaining higher accuracy.

[0039] Specifically, after the model is input, it needs to pass through a convolutional layer of the MCC module. The convolutional kernel size of this convolutional layer is 3, and in order to keep the shape of the feature layer unchanged, it is padded with 1 around it. Specifically, the relational expression of the MCC module is: MCC = BN(GeLu(Conv(BachNormal(GeLu(Mamba(x)))))). Further, after being processed by the convolutional layer, in this example, the GeLu activation function is used to perform nonlinear rectification on the features. The GeLu activation function can effectively improve the robustness of the model and enable the model to improve the generalization ability on the training samples. Thus, through nonlinear rectification, a weight layer for discriminating the lodged image is obtained. However, considering that with the increase of the number of network layers, the problems of gradient explosion or gradient disappearance are likely to occur. Therefore, in this embodiment, see Figure 2 , a normalized BatchNorm2D layer is added to the MCC module.

[0040] In the long-term study of wheat lodging, the inventor found that the traditional use of convolutional neural network models is limited by the size of the receptive field, and the size of its convolutional kernel determines that it cannot perform convolutional analysis on images, so it is impossible to extract the visual difference image features between lodged wheat and unlodged wheat. For this reason, in this embodiment, the SSMoudle module is introduced into the UssNet model. It effectively calculates the correlation of the same category in the image based on Visual Mamba through MLP to achieve image classification. And MLP does not use a convolutional kernel, and the perspective of its receptive field is on each pixel of the image, so it can better match the same category. Specifically, see Figure 2 , the input features first encounter a linear embedding layer and then bifurcate into a dual path. One branch passes through depth convolution and SiLU activation and enters the SSM module with strong spatial and context awareness capabilities. After layer normalization, it is merged with the alternate stream after SiLU activation and output.

[0041] In this embodiment, the relational expression of the CCR module is specifically CCR=(GeLu(Conv(BachNormal(GeLu(Conv(x)))))). Its double-layer convolution characteristic enables the model to capture more feature information under the same feature dimension. And to avoid the influence of dead neurons of the activation function on the generalization and convergence stability of the model, this module uses the GeLu activation function. Similarly, considering that with the increase of the number of network layers, problems such as gradient explosion or gradient disappearance are likely to occur. For this reason, in this embodiment, see Figure 2 , a normalized BatchNorm2D layer is also added to the CCR module. Thus, a convolutional layer, a GeLu activation function, a normalization layer, and a convolutional layer are combined through an information integration module to further improve the model's ability to extract image textures and graphics.

[0042] In addition, to alleviate the forgetting of shallow features by deep layers in the model, see Figure 2 , in this embodiment, a residual module RestBlock is also added as a residual layer and introduced into the MCC module and the CCR module respectively.

[0043] (4) Adopt the Focal loss function and train the UssNet model based on the sample dataset.

[0044] Since in the obtained UAV images, the area gap between the wheat lodging area and the normal growth area is large, the proportion of normal samples and lodging samples in the finally obtained sample dataset is extremely uneven. In order to suppress the influence of sample imbalance on the calculation of the loss value, in this embodiment, the Focal loss function is adopted. The Focal loss function can enhance the model's attention to small targets and improve the small target recognition ability. Specifically, the Focal loss function includes a modulation factor (1-fk ) α , where there is an adjustable focus parameter α (α > 0) to reduce the model's attention to simple samples, so that the model focuses on the lodging area. Among them, the expression of the Focal loss function is: .

[0045] In addition, in this embodiment, when training the UssNet model based on the sample dataset, the average running time of the lodging sample dataset is defined to include data transmission, model training, and inference processes. The sample dataset is specifically processed by 2D image segmentation. To avoid overfitting, the dropout random discard algorithm is adopted, and the hyperparameter is set to 0.1. The training epoch of the wheat lodging model is set to 300 rounds, and the batch size is 4. The initial learning rate is 0.001, the learning rate update strategy adopts the cosine annealing strategy, the weight decay coefficient is set to 0.001, and the optimization algorithm adopts AdamW.

[0046] To evaluate the reliability and accuracy of the UssNet model, in this embodiment, precision (Pre), recall (Rec), F1 score , and IoU are used to evaluate each category, and nIoU is used to test the entire model. The values of the above indicators range from 0 to 1, and the larger the value, the better the evaluation effect. The calculation formulas of the evaluation indicators are as follows: Pre = TP / (TP + FP).

[0047] Rec = TP / (TP + FN).

[0048] F1 score = 2TP / (2TP + FP + FN).

[0049] IoU = TP / (TP + FP + FN).

[0050] .

[0051] Among them, TP represents true positive, FP represents false positive, FN represents false negative, and TN represents true negative.

[0052] To verify the effectiveness and reliability of the wheat lodging area recognition model of the present disclosure, in this embodiment, Unet, SegNet, DeepLabv3+, and ViT models are also established respectively, and the recognition performances of each model are compared. Among them, the UNet model can be well trained with fewer samples; the SegNet model performs image segmentation through a symmetric encoding and decoding structure, uses pooling indices to save pooling position information, can retain more details and obtain better segmentation results; the DeepLabv3+ network adds an encoding and decoding module and is improved from the Xception backbone network, and uses dilated convolution and depthwise separable convolution to improve the segmentation accuracy; while the ViT model is based on the ViT architecture and uses the self-attention mechanism to capture global information by treating the image as a sequence of continuous patches, which enhances the model's understanding of long-range dependencies and thus improves the segmentation accuracy.

[0053] To control the variables of the comparative experiment and ensure the fairness of model comparison, in this embodiment, all experiments are carried out on the same Ubuntu 18.04 system. The code structures of all computing programs are the same, using Python 3.10, PyTorch 2.1, and CUDA 11.8. And the training parameters of all wheat lodging models are set to be the same.

[0054] In addition, since the originally collected images are large, when inputting images into the network, the sizes of the annotated images and the predicted images are uniformly adjusted to 512×512 pixels. Randomly select 10 wheat lodging images, and automatically classify the trained model on the actually obtained lodging images. See Figure 3 , the green in the prediction map is judged as lodging, and the black is judged as non-lodging. From Figure 3 It can be seen that the verification results of UNet, ViT (Transformer), SegNet, and DeepLabv3+ are poor. In the wheat lodging segmentation task, these 5 network models cannot well distinguish between lodging and non-lodging wheat. For example, for the UNet model, see Figure 3 In the lower area of the fourth picture in the fourth row of

[0055] Figure 4 shows the performances of each model on the complete remote sensing images of wheat lodging. From Figure 4It is not difficult to see that among the five models, UssNet shows better performance in details such as scattered lodging areas of wheat. In the present invention, UssNet adopts the Mamba structure to well make up for the shortcomings of other models. And by adding the SSMoudle module, this module has powerful spatial and context awareness capabilities, and its receptive field perspective is on each pixel of the image, which can well match the lodging area. In addition, the UssNet network model comprehensively considers the characteristics of lodging wheat images and the data distribution characteristics of lodging areas, and the detection of the boundary of lodging wheat is most consistent with the ground truth. In contrast, SegNet, DeepLabv3+ and ViT show classification errors at the boundaries of regions. Although these two categories have obvious and highly distinguishable features, there are some small background regions and blurred boundaries between normal categories, making it difficult for these models to accurately distinguish. The good segmentation effect of the UssNet network model proves, on the one hand, the importance of adding global perception in convolutional neural network models; on the other hand, it proves that simply increasing the scale of the number of parameters and the complexity of the model in convolutional neural networks does not necessarily improve the segmentation accuracy. Instead, it is necessary to perceive global features in order to improve the segmentation fineness and accuracy.

[0056] See Table 1. Table 1 shows the performance metrics of UssNet and other semantic segmentation models on the wheat lodging dataset, where the parameter unit is in millions (M). It can be seen from Table 1 that the UssNet model has obvious advantages in terms of the number of parameters and computational complexity. In terms of performance, UssNet achieves the highest PA (0.969) and mIoU (0.928) among all models, indicating a good balance between performance and resource efficiency. Compared with UNet, SegNet, DeepLabV3+, and ViT, UssNet has a higher recognition accuracy in 4 evaluation metrics.

[0057] 。

[0058] Figure 5 It shows the change of mIoU of the wheat lodging area with the number of epochs and the pixel accuracy of each model. As Figure 5 (a) shows, in the first few rounds of iteration of each model, the mIoU of the wheat lodging area rises sharply and the loss value drops rapidly. As the number of rounds increases, the accuracy and loss value of the network remain stable. Due to the existence of the learning rate adjustment strategy, the learning rate of each model gradually decays with the increase of the number of iterations. As Figure 5As shown in (a - b), after 100 rounds of training, although all curves fluctuate slightly, the wheat lodging mIoU and overall accuracy of each model remain stable. Among them, the mIoU of the wheat lodging category and the overall pixel accuracy of UssNet are the highest, reaching 0.928 and 0.969 respectively. Compared with the UNet, SegNet, DeepLabv3 +, and ViT models, the mIoU of the UssNet model is 10.02%, 11.9%, 9.61%, and 41.4% higher respectively. The pixel accuracy value of the UssNet model is 6.7%, 7.3%, 8.5%, and 16.9% higher than that of the above network models respectively. Generally speaking, UssNet has better and more stable segmentation performance than other network models in the recognition of long - distance dependencies in lodging images and regions with small lodging areas.

[0059] In addition, refer to Figure 6 , and select 5 photos to compare the recognition effects of Grad - CAM and 3D weight mapping visualization techniques in 2D and 3D lodging edge regions. From Figure 6 The results shown, it can be clearly seen that when the model classifies wheat lodging, it mainly focuses on the boundary regions in the image. For example, in Figure 6 , the red area represents the region with higher model attention, which indicates that when the model makes a decision, it mainly relies on the feature information of these regions. The key regions shown by Grad - CAM and 3D weight mapping are somewhat consistent with our expected judgment basis for boundary regions, further verifying the rationality of the UssNet model's decision - making. At the same time, by comparing the visualization results of different images, it can also be found that the UssNet model has obvious distinguishability for the key feature regions of boundary information, which helps to better understand the classification mechanism of the model and the feature differences between different categories.

[0060] In this example, considering that when actual wheat lodging disasters occur, the lodging regions have diverse characteristics, and the shape and size of each wheat lodging region will vary, resulting in a large difference in the proportion of different wheat lodging regions in the acquired UAV images, and finally there is a problem of uneven distribution of lodging samples in the constructed UAV image dataset. In order to improve the impact of uneven sample distribution on the calculation of the loss function, this example compares two loss functions: the cross - entropy loss function and the Focal loss function used in this embodiment. Figure 7 and 8 show the comparison results of using different loss functions in the UssNet, UNet, SegNet, DeepLabv3 +, and ViT models. From Figure 7 and 8It can be seen that in the comparison of loss functions, the Focal loss function performs well in the 5 models. In terms of the overall evaluation index of the model, the comprehensive evaluation index of UssNet is the highest. In the Focal loss function, PA is 0.9693 and mIou is 0.878. This is related to the fact that the Focal loss function balances the loss between easy and hard samples by reducing the contribution of easy-to-classify samples. At the same time, some fluctuations are observed from the curve trend because the model is too sensitive to the data. In this example, 5052 datasets are used, which will cause insufficient samples to a certain extent, thus causing the oscillation phenomenon of the model training accuracy.

[0061] To analyze the effectiveness of each component of UssNet, an ablation study was conducted on each component of the network model. The ablation experiment is an experiment designed to determine the impact of each part on the experimental results. The results are as Figure 9 shown. It can be clearly seen from the results that when we use SSM to replace the Conv block, the reconstruction results show varying degrees of degradation. As Figure 9 (b) and (c) show, the feature layer with the addition of the state space model can extract more abundant information because the images obtained from remote sensing have height information, and this feature can also be utilized in the task of lodging detection to improve the accuracy. There is a height difference between the lodged wheat and the unlodged wheat. However, the conventional convolutional kernel is limited by the range of its local receptive field and actually cannot perceive the entire image, while the addition of the spatial perception module can significantly improve the effect of spatial perception.

[0062] To verify the effectiveness of the SSMoudle module, activation function, and loss function in UssNet, this example conducts ablation experiments by gradually adding modules to analyze their impact on model performance. The experimental results are shown in Table 2. Among them, UNet represents the baseline model, which uses ordinary convolutions to build the backbone network, the ReLu activation function, and does not contain the SSMoudle module. UNet_Loss represents adding a contrastive loss function to the baseline model, UssNet_GeluLoss represents adding the GeLu activation function and contrastive loss function to the baseline model, and UssNet_SSMGeluLoss represents adding the SSMoudle module, GeLu activation function, and loss function to the baseline model. From the ablation experiment results in Table 2, it can be seen that the modules all effectively improve the model performance. Specifically, when the SSMoudle module and GeLu activation function are added to the baseline model, the accuracy of the model can be improved by about 4.1 percentage points, Recall is improved by about 4 percentage points, and F1Score is improved by about 3.5 percentage points. It is worth noting that when the SSMoudle module and activation function are added to the model, the accuracy of the model in PA is significantly improved, from 92.3% to 94.9%. This result further emphasizes that the UssNet model pays more attention to the wheat lodging area and global recognition ability, and suppresses irrelevant information.

[0063] 。

[0064] Although some preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0065] Obviously, those skilled in the art can make various changes and variations to the present disclosure without departing from the spirit and scope of its inventive concept. Thus, if these modifications and variations to the present disclosure fall within the scope of the claims of this application and their equivalent technologies, this application also intends to include these modifications and variations.

Claims

1. A method for establishing a wheat lodging area recognition model, characterized in that The method includes the following steps: (1) Obtain wheat lodging images by a multispectral drone under clear and cloudless weather conditions; (2) Preprocess the wheat lodging images obtained by the multispectral drone to obtain a sample data set; And correspondingly divide the data set into a training set, a validation set and a test set; (3) Establish a UssNet model including an MCC module, a CCR module, an SSModule module and a RestBlock module; the encoder of the UssNet model is correspondingly composed of the MCC module and the CCR module in a cross manner, and the decoder of the UssNet model corresponds to the structure of the encoder; (4) Use the Focal loss function and train the UssNet model based on the sample data set.

2. The method for establishing a wheat lodging area recognition model according to claim 1, wherein In the step (2), the wheat lodging images with inclined fields of view and irrelevant shadows are removed in advance before the preprocessing; the preprocessing includes image cropping for cropping the wheat lodging images by a rectangular window of a specific size to generate the images of the sample data set and image annotation for annotating the contours of the wheat lodging areas.

3. The method for establishing a wheat lodging area recognition model according to claim 1, characterized in that, In the step (2), perform angle transformation and data augmentation processing on the images belonging to the wheat lodging category in the sample data set, and correspondingly divide all the processed extended data into the training set.

4. The method for establishing a wheat lodging area recognition model according to claim 1, characterized in that In the step (3), the MCC module performs convolution processing on the model input and then uses the GeLu activation function to perform non-linear rectification on the features; and the MCC module also includes a normalized BatchNorm2D layer.

5. The method for establishing a wheat lodging area recognition model according to claim 4, wherein, In the step (3), the SSModule module uses the features after non-linear rectification by the GeLu activation function as the input, and uses a dual-path processing data respectively including a linear embedding layer as the output, and sequentially performs depth convolution, SiLu activation, SSM module and layer normalization processing along one path branch.

6. The method for establishing a wheat lodging area recognition model according to claim 1, characterized in that In the step (3), the CCR module includes two convolutional layers and uses the GeLu activation function.

7. The method for establishing a wheat lodging area recognition model according to claim 1, wherein In the step (3), the RestBlock module is correspondingly introduced into the MCC module and the CCR module respectively.

8. The method for establishing a wheat lodging area recognition model according to claim 1, characterized in that In the step (4), the model training uses the dropout random discard algorithm, the hyperparameter is set to 0.1, the training epoch is 300 rounds, the batch is 4, the initial learning rate is 0.001, the learning rate update strategy is the cosine annealing strategy, the weight decay coefficient is 0.001, and the optimization algorithm is AdamW.

9. A method for identifying wheat lodging areas, which identifies wheat lodging areas based on the UssNet model trained correspondingly according to Claim 1.