Grassland restoration survival rate evaluation method and system
By improving the CS-YOLOv7 and MS-DeepLabv3+ networks, the accuracy and efficiency issues of grassland seedling detection and segmentation are solved, enabling efficient and accurate assessment of grassland restoration survival rate, which is suitable for UAV edge computing devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CCCC FIRST HIGHWAY ENG GRP HUAZHONG ENG CO LTD
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-28
AI Technical Summary
Traditional manual assessments and existing automated solutions cannot accurately capture the small target features of grassland seedlings, leading to positioning errors. Existing semantic segmentation models have a large number of parameters and high computational complexity, making them difficult to adapt to edge computing devices such as drones. This affects the efficiency and accuracy of grassland seedling survival rate assessment and makes it difficult to meet the real-time monitoring and accurate assessment needs of large-scale grassland ecological restoration.
An improved CS-YOLOv7 object detection network and MS-DeepLabv3+ semantic segmentation network were constructed. YOLOv7 was optimized by inserting convolutional block attention modules, simple attention modules, bilinear interpolation, and Mish activation functions. DeepLabv3+ was optimized by combining lightweight MobileNetv3 and squeeze-excited attention modules to achieve accurate detection and segmentation of seedlings in grassland infrared images.
The improved network model can accurately capture the small target features of seedlings, improve the accuracy of location positioning and segmentation, reduce computational requirements, adapt to UAV equipment, and improve the efficiency and accuracy of grassland seedling survival rate assessment.
Smart Images

Figure CN121937835A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of grassland ecological monitoring technology, and in particular to a method and system for assessing the survival rate of grassland restoration. Background Technology
[0002] Grasslands, as an important component of terrestrial ecosystems, play crucial ecological roles in soil and water conservation, climate regulation, and biodiversity maintenance. However, due to factors such as overgrazing, climate change, and human-caused damage, grasslands in many parts of the world are degrading, making grassland restoration one of the core tasks of ecological environmental protection. The grassland restoration survival rate, as a core indicator for evaluating the effectiveness of restoration projects, is crucial for accurate and rapid assessment, which is essential for adjusting subsequent restoration strategies and optimizing resource allocation.
[0003] Currently, deep learning-based grassland seedling detection and segmentation technologies mainly revolve around object detection networks and semantic segmentation networks. In the field of object detection, models such as YOLOv7 and Faster R-CNN are widely used. Among them, YOLOv7 has become the preferred solution for real-time detection scenarios due to its high efficiency in single-stage detection, achieving feature extraction and bounding box regression through the LeakyReLU activation function and CIoU loss function. In the field of semantic segmentation, DeepLabv3+ is the mainstream model, which uses Xception as the encoder backbone network and combines the multi-scale dilated convolution of the ASPP module with feature fusion of the decoder to achieve pixel-level semantic classification.
[0004] However, traditional manual assessment and existing automated solutions cannot accurately capture the small target features of seedlings, which can easily lead to positioning errors. Existing semantic segmentation models are also difficult to adapt to edge computing devices such as drones due to their large number of parameters and high computational complexity. Moreover, the models are not adaptable enough to the feature extraction of seedlings at different scales, resulting in poor segmentation results. This seriously affects the efficiency and accuracy of grassland seedling survival rate assessment and makes it difficult to meet the real-time monitoring and accurate assessment needs of large-scale grassland ecological restoration. Summary of the Invention
[0005] To address the challenges of traditional manual assessments and existing automated solutions failing to accurately capture the small features of seedlings, which can easily lead to positioning errors, and the fact that existing semantic segmentation models suffer from large parameter counts and high computational complexity, making them difficult to adapt to edge computing devices such as drones; furthermore, the models lack adaptability to feature extraction from seedlings of different scales, resulting in poor segmentation performance. These issues severely impact the efficiency and accuracy of grassland seedling survival rate assessment, making it difficult to meet the technical requirements for real-time monitoring and accurate assessment of large-scale grassland ecological restoration. This invention provides a method and system for assessing grassland restoration survival rate.
[0006] The technical solutions provided by the embodiments of the present invention are as follows: The first aspect of this invention provides a method for assessing the survival rate of grassland restoration, comprising: S1: Collect sample data of infrared images of grassland; S2: Preprocess the grassland infrared image sample data to obtain standard grassland infrared image sample data; S3: Construct a CS-YOLOv7 object detection network based on the original YOLOv7; S4: Construct the MS-DeepLabv3+ semantic segmentation network, which is an improvement on the original DeepLabv3+; S5: Input the labeled data from the standard grassland infrared image sample data into the CS-YOLOv7 target detection network and the MS-DeepLabv3+ semantic segmentation network respectively for training to obtain the seedling target detection model and the semantic segmentation model. S6: Acquire infrared image data of the grassland to be detected; S7: Input the infrared image data of the grassland to be detected into the seedling target detection model, and output the seedling location information and seedling survival status; S8: Segment the seedling location information using a semantic segmentation model; S9: Calculate the survival rate of grassland seedlings based on the detection and segmentation results.
[0007] A second aspect of the present invention provides a grassland restoration survival rate assessment system, comprising: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the grassland restoration survival rate assessment method as described in the first aspect.
[0008] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the grassland restoration survival rate assessment method as described in the first aspect.
[0009] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, the improved CS-YOLOv7 target detection network can accurately capture the small target features of seedlings in infrared images, effectively improving the localization accuracy of seedling positions. The improved MS-DeepLabv3+ semantic segmentation network significantly reduces the model's computational requirements while ensuring segmentation accuracy, enabling efficient deployment on mobile devices such as drones. Simultaneously, segmenting seedling position information using the semantic segmentation model can cover the scale range of seedlings at different growth stages, improving the full-scale target segmentation accuracy from seedling to mature plant, ensuring the accuracy of seedling region division, and ultimately improving the efficiency and accuracy of grassland seedling survival rate assessment. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart illustrating a grassland restoration survival rate assessment method provided in an embodiment of the present invention.
[0012] Figure 2 This is a diagram of a CS-YOLOv7 target detection network architecture provided for an embodiment of the present invention.
[0013] Figure 3 This is a diagram of an MS-DeepLabv3+ semantic segmentation network architecture provided for an embodiment of the present invention.
[0014] Figure 4 This is a schematic diagram of a grassland restoration survival rate assessment system provided in an embodiment of the present invention. Detailed Implementation
[0015] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0016] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0017] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, their intended meanings are consistent. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, their intended meanings are consistent.
[0018] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0019] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0020] Reference manual attached Figure 1 The diagram shows a flowchart of a grassland restoration survival rate assessment method provided by an embodiment of the present invention.
[0021] This invention provides a method for assessing the survival rate of grassland restoration projects. This method can be implemented using a grassland restoration survival rate assessment device, which can be a terminal or a server. The processing flow of the grassland restoration survival rate assessment method may include the following steps: S1: Collect sample data of infrared images of grassland.
[0022] Optionally, the grassland infrared image sample data specifically includes: temperature resolution, thermal sensitivity, image resolution, acquisition frame rate, emissivity, and object distance.
[0023] It should be noted that a hexacopter UAV was selected as the experimental platform, integrating a DY640 infrared thermal imaging system for core data acquisition, and simultaneously configured with a three-dimensional laser ranging module (sampling at the same frequency as the thermal imaging system) and a high-definition image real-time transmission device. The specific data for the grassland infrared image samples are as follows: temperature resolution 0.04℃, thermal sensitivity 0.035℃, image resolution 640×480, acquisition frame rate 25FPS, emissivity 0.7-0.75, and object distance 4m-6m.
[0024] Furthermore, during the data collection process, the UAV is controlled to maintain a flight altitude of 10-15m and a flight speed of 5m / s. It conducts full-coverage image collection of the target forest and grassland restoration area (such as degraded grassland restoration area and mountain afforestation area) according to the preset route to ensure complete image coverage and no obvious blurring. At the same time, the collection parameters are recorded, including flight altitude, shooting angle, ambient temperature, laser ranging data, etc.
[0025] S2: Preprocess the grassland infrared image sample data to obtain standard grassland infrared image sample data.
[0026] In one possible implementation, S2 specifically includes sub-steps S201 to S204: S201: Temperature calibration is performed on the image pixel values in the grassland infrared image sample data based on the ambient temperature parameters.
[0027] S202: Denoising is performed on temperature-calibrated grassland infrared image sample data by combining Gaussian filtering and median filtering.
[0028] Gaussian filtering is a linear smoothing filtering method based on the Gaussian function. It assigns different weights to the target pixel and all pixels in its neighborhood in an image (pixels closer to the target pixel have higher weights, and vice versa), and then calculates the new gray value of the target pixel through a weighted average. First, a corresponding Gaussian kernel matrix is generated according to the preset filtering window size and the standard deviation of the Gaussian kernel. The numerical distribution of the matrix follows a normal distribution. Then, the Gaussian kernel is slid across the image pixel by pixel, and the weight of each position in the kernel is multiplied by the gray value of the corresponding pixel and then summed. The final result is the gray value of the target pixel after filtering.
[0029] Median filtering is a nonlinear signal processing and image denoising method based on sorting statistics. It replaces the gray value of a pixel in an image or signal with the median, rather than the mean, of all the gray values of pixels within its neighborhood window. First, a fixed-size sliding window is set. All pixel gray values within the window's coverage area are sorted by size. The value at the middle of the sorted sequence is taken as the new gray value for the current pixel. Then, the window is slid pixel by pixel to complete the filtering process for the entire image.
[0030] S203: Enhance image contrast of denoised grassland infrared image sample data through histogram equalization.
[0031] Histogram equalization is a classic image contrast enhancement technique. By adjusting the gray-level histogram distribution of an image, it stretches the concentrated gray-level range in the original image to a more uniform distribution range, thereby improving the overall contrast of the image. First, the number of pixels at each gray level in the image is counted and the cumulative distribution function is calculated. Then, a mapping transformation is used to convert the original gray-level values into new gray-level values, making the gray-level distribution of the processed image closer to a uniform state.
[0032] S204: Label the grassland infrared image sample data after image contrast enhancement to obtain standard grassland infrared image sample data.
[0033] Optionally, the annotation content of the standard grassland infrared image sample data specifically includes: the bounding box of the seedling area, the seedling outline, and the survival status label.
[0034] In this embodiment of the invention, temperature calibration is used to eliminate the interference of ambient temperature on the pixel values of grassland infrared images, ensuring the authenticity and consistency of image data. A combined denoising method of Gaussian filtering and median filtering is adopted, which utilizes the linear smoothing characteristics of Gaussian filtering to weaken Gaussian noise in the image, and takes advantage of the nonlinear processing of median filtering to remove impulse noise, effectively preserving image edge and detail information. Histogram equalization is used to stretch the dynamic range of image grayscale, enhancing the contrast between seedlings and background and improving the recognition of target areas. Finally, the image is accurately labeled to generate standard sample data containing seedling bounding boxes, contours, and survival status labels. This not only provides high-quality and reliable input data for subsequent model training, but also effectively improves the model's detection accuracy and classification efficiency for grassland seedlings, providing accurate data support for grassland ecological monitoring and restoration assessment.
[0035] Reference manual attached Figure 2 The diagram illustrates a CS-YOLOv7 target detection network architecture provided by an embodiment of the present invention.
[0036] It should be noted that the instruction manual includes... Figure 2 In this diagram, INPUT represents the input image, Backbone represents the backbone network, CBM represents the convolution-batch normalization-activation module, CB5A1 represents the improved convolution-batch normalization-activation module, MP represents the max pooling layer, CS5A represents the first feature extraction submodule, CB5B represents the improved convolution-batch normalization-activation module, CS5B represents the second feature extraction submodule, Head represents the detection head, CBC represents the optimized convolution-batch normalization-activation module, UpSample represents the upsampling layer, Concat represents the concatenation layer, and C5 represents the feature map output layer.
[0037] The backbone network specifically includes: a first convolution-batch normalization-activation module, a second convolution-batch normalization-activation module, a first convolution-batch normalization-activation improved module, a first max pooling layer, a first feature extraction submodule, a second max pooling layer, a second convolution-batch normalization-activation improved module, a third max pooling layer, and a second feature extraction submodule.
[0038] Furthermore, in the backbone network, the first convolution-batch normalization-activation module, the second convolution-batch normalization-activation module, the first convolution-batch normalization-activation improved module, the first max pooling layer, the first feature extraction submodule, the second max pooling layer, the second convolution-batch normalization-activation improved module, the third max pooling layer, and the second feature extraction submodule are connected sequentially from left to right.
[0039] The detection head specifically includes: a first convolutional-batch normalization-activation optimization module, a third convolutional-batch normalization-activation module, a first upsampling layer, a fourth convolutional-batch normalization-activation module, a first stitching layer, a first feature map output layer, a fifth convolutional-batch normalization-activation module, a second upsampling layer, a sixth convolutional-batch normalization-activation module, a second stitching layer, a second feature map output layer, a seventh convolutional-batch normalization-activation module, an eighth convolutional-batch normalization-activation module, a ninth convolutional-batch normalization-activation module, a third stitching layer, a third feature map output layer, a tenth convolutional-batch normalization-activation module, an eleventh convolutional-batch normalization-activation module, a twelfth convolutional-batch normalization-activation module, a fourth stitching layer, a fourth feature map output layer, a thirteenth convolutional-batch normalization-activation module, and a fourteenth convolutional-batch normalization-activation module.
[0040] Furthermore, the second feature extraction submodule in the backbone network is connected to the first convolutional-batch normalization-activation optimization module. The first convolutional-batch normalization-activation optimization module is sequentially connected to the third convolutional-batch normalization-activation module and the first upsampling layer. The output of the second convolutional-batch normalization-activation improvement module in the backbone network is connected to the fourth convolutional-batch normalization-activation module. The fourth convolutional-batch normalization-activation module and the first upsampling layer are sequentially connected to the first concatenation layer, the first feature map output layer, the fifth convolutional-batch normalization-activation module, and the second upsampling layer. The output of the first feature extraction submodule in the backbone network is connected to the sixth convolutional-batch normalization-activation module. The sixth convolutional-batch normalization-activation module and the second upsampling layer are sequentially connected to the second concatenation layer and the second feature map output layer. The eighth convolutional-batch normalization-activation module and the seventh convolutional-batch normalization-activation module output H3. The second feature map output layer is connected to the ninth convolutional-batch normalization-activation module. The outputs of the ninth convolutional-batch normalization-activation module and the first feature map output layer are sequentially connected to the third concatenation layer, the third feature map output layer, the eleventh convolutional-batch normalization-activation module, and the tenth convolutional-batch normalization-activation module, outputting H4. The third feature map output layer is connected to the twelfth convolutional-batch normalization-activation module. The outputs of the twelfth convolutional-batch normalization-activation module and the first convolutional-batch normalization-activation optimization module are sequentially connected to the fourth concatenation layer, the fourth feature map output layer, the thirteenth convolutional-batch normalization-activation module, and the fourteenth convolutional-batch normalization-activation module, outputting H5.
[0041] S3: Construct a CS-YOLOv7 object detection network based on the original YOLOv7.
[0042] YOLOv7 consists of three parts: a backbone network, a neck network, and a head. The backbone network uses stacked Efficient Layer Aggregation Network (ELAN) modules to extract deep features and enhances the model's non-linear expressive ability through multi-level convolution and activation functions. The neck network completes multi-scale feature fusion based on the Path Aggregation Network (PAN) structure, upsampling, downsampling, and concatenating the features from different levels output by the backbone network to combine shallow detail features with deep semantic features. The head outputs the target's location, category, and confidence score through anchor box regression and classification branches. It also introduces the CIoU loss function to optimize the bounding box regression accuracy and adopts a dynamic label allocation strategy to improve the detection performance of small targets.
[0043] The CS-YOLOv7 object detection network is an improved object detection network based on YOLOv7. Its core consists of two parts: a backbone network and a head. The backbone network completes multi-scale abstraction and extraction of input features from shallow to deep layers through the concatenation of CBM (convolution-batch normalization-activation module), CB5A / CB5B (improved convolution module), MP (max pooling layer), and CS5A / CS5B (feature extraction sub-module). The head integrates components such as convolution optimization module, upsampling layer, and concatenation layer. Through the multi-level connection logic of "convolution-upsampling-feature concatenation", it fuses and refines the features of different levels output by the backbone network, and finally outputs the detection results.
[0044] In one possible implementation, S3 specifically includes sub-steps S301 to S305: S301: Insert a convolutional block attention module after the original YOLOv7 C3 module.
[0045] The C3 module is a core lightweight feature extraction module in the backbone of the YOLO series object detection network. Based on residual learning and cross-layer connections, its overall structure includes a main branch and a residual branch. The main branch achieves deep feature extraction by stacking multiple basic convolutional modules (such as CBM, which consists of convolution, batch normalization, and activation functions). The residual branch adjusts the feature dimensions through a small number of convolutional layers and then fuses them with the output features of the main branch. At the same time, the module adopts a bottleneck structure to reduce the number of parameters and computational complexity.
[0046] Among them, the Convolutional Block Attention (CBAM) module is a lightweight channel-space dual attention mechanism module. Through the "channel attention branch" (which weights the importance of feature channels) and the "spatial attention branch" (which weights the importance of feature spatial regions), the model can automatically focus on key feature regions and suppress irrelevant noise. It is often inserted into network modules to enhance the ability to extract target features.
[0047] S302: Insert a simple attention module after the fast spatial pyramid pooling module in the original YOLOv7.
[0048] The Fast Spatial Pyramid Pooling (SPPF) module is a lightweight multi-scale feature extraction module in the backbone of the YOLO series networks. It is an optimized upgrade of the traditional Spatial Pyramid Pooling (SPP) module. It captures features of different target sizes through multi-scale pooling operations, solving the problem of insufficient feature extraction caused by differences in target scale in target detection. The input feature map is fed into the same pooling window for max pooling, and then the pooled multi-scale features are concatenated and fused with the original input features, ultimately outputting a feature map containing rich scale information. Compared to the traditional SPP module, the SPPF module significantly reduces computational complexity and inference time while maintaining multi-scale feature extraction capabilities by sharing convolutional kernels and simplifying the computational process, making it more suitable for real-time target detection scenarios.
[0049] The SimAg Attention (SimAM) module is a parameter-free, lightweight neuron-level attention mechanism that enhances key features and suppresses irrelevant background noise by calculating the importance weight of each neuron in the feature map. Unlike channel attention or spatial attention modules that require additional convolutional layers to learn weights, the SimAg module is directly based on the neuroscience principle of "neurons maximizing their own signal-to-noise ratio." It derives and calculates the dependency relationship between each neuron and its neighboring neurons to generate corresponding attention weights, which are then applied to the original feature map to highlight the feature response of the target region. Without introducing additional parameters or computational overhead, it can be flexibly embedded into any feature layer of various convolutional neural networks. In object detection networks such as YOLOv7, deploying it after the SPPF module can accurately enhance the feature signals of small targets like seedlings, improving the model's recognition accuracy for seedlings of different scales, while avoiding the computational burden caused by adding modules.
[0050] S303: Replace the detection head portion in the original YOLOv7 with bilinear interpolation.
[0051] The head is the core output module of the object detection network, located after the backbone and neck network. It is responsible for transforming the multi-scale feature maps extracted and fused by the preceding modules into the final object detection result. In single-stage detection algorithms such as YOLOv7, the head typically includes a classification branch and a regression branch: the classification branch calculates the confidence of the object category to which each predicted box belongs through convolution and activation functions, while the regression branch is responsible for predicting the coordinate parameters of the object bounding box, and simultaneously optimizes the localization accuracy of the predicted box by combining anchor box mechanism and loss function.
[0052] Bilinear interpolation is a classic nonlinear interpolation algorithm used to scale images or feature maps (upsampling or downsampling). It calculates the final value of a target pixel by weighted averaging the grayscale values (or feature values) of its four neighboring pixels. First, linear interpolation is performed on two adjacent pixels horizontally, and then a second linear interpolation is performed on the interpolation result vertically. The resulting interpolation is smooth and continuous, effectively avoiding the jagged distortion caused by neighbor interpolation. In object detection networks, replacing the detection head with a bilinear interpolation module enables fine-grained scaling of feature maps, improving the boundary localization accuracy of small targets (such as seedlings in grasslands). Furthermore, this algorithm has low computational complexity, does not significantly increase the model's inference burden, and is suitable for deployment requirements of edge devices such as drones.
[0053] It should be noted that interpolation calculations are performed by considering the relative distances and weights of multiple pixel values around the target location to improve the positioning accuracy of small targets.
[0054] S304: Replace the LeakyReLU activation function in the original YOLOv7 with the Mish function.
[0055] LeakyReLU activation function is an improved nonlinear activation function compared to the traditional ReLU activation function, used to address the "neuron death" problem caused by the gradient being zero in the negative region of ReLU. In object detection networks such as YOLOv7, LeakyReLU is often embedded in convolutional modules to enhance the network's nonlinear expressive power, improve the model's fitting accuracy for complex features, and avoid training stagnation caused by neuron inactivation, thus ensuring the stable convergence of deep networks.
[0056] The Mish function is a smooth, non-monotonic, non-linear activation function. By combining the saturation properties of the tanh function with the smoothing properties of the softplus function, it achieves continuous differentiability of the output curve. Compared to piecewise activation functions such as ReLU and LeakyReLU, the Mish function does not exhibit hard saturation or gradient vanishing in the negative interval, providing a more stable gradient flow for training deep networks. In object detection networks such as YOLOv7, the Mish function is often used in the feature extraction module of the backbone network. With its excellent feature representation capabilities, it can effectively improve the model's accuracy in capturing features of small targets (such as seedlings in grass) in complex scenes. Simultaneously, the smoothness of this function helps reduce oscillations during training and accelerates model convergence.
[0057] It should be noted that the superior adaptability and smoothness of the Mish function allow it to better adapt to different data distributions and model complexities, thereby improving nonlinear modeling capabilities and convergence speed.
[0058] S305: Construct a CS-YOLOv7 object detection network based on the C3 module, simple attention module, bilinear interpolation, and Mish function.
[0059] Furthermore, the CS-YOLOv7 object detection network uses the DIoU loss function to optimize bounding box regression, reduces the optimal loss value by redefining the penalty metric, and improves the bounding box localization accuracy. At the same time, it combines the cross-entropy loss function to complete the seedling survival status classification.
[0060] In this embodiment of the invention, a multi-dimensional performance optimization is achieved by constructing a CS-YOLOv7 target detection network based on an improvement of the original YOLOv7. A Convolutional Block Attention (CBAM) module is inserted after the C3 module, using a dual channel and spatial attention mechanism to allow the model to automatically focus on key target feature regions such as grass seedlings, suppressing background noise interference. A Simple Attention (SimAM) module is embedded after the Fast Spatial Pyramid Pooling (SPPF) module, using a parameter-free and lightweight design to enhance the signal response of seedling targets in multi-scale features, improving the accuracy of small target recognition without increasing computational burden. The detection head is replaced with a bilinear interpolation module, optimizing the seedling boundary localization effect through refined feature map scale adjustment, avoiding interpolation distortion. The LeakyReLU activation function is replaced with the Mish function, utilizing its smooth and continuous nonlinear characteristics to improve gradient flow in deep networks, accelerate model convergence, and enhance feature fitting capabilities in complex scenes. Finally, by combining the DIoU loss function to optimize the bounding box regression accuracy and the cross-entropy loss function to complete the seedling survival status classification, the constructed CS-YOLOv7 network balances detection accuracy and inference efficiency, and can accurately adapt to the detection needs of seedling targets in grassland infrared images, providing reliable technical support for grassland ecological monitoring.
[0061] Reference manual attached Figure 3 The diagram illustrates an MS-DeepLabv3+ semantic segmentation network architecture provided by an embodiment of the present invention.
[0062] It should be noted that the instruction manual includes... Figure 3 In this context, INPUT represents the input image, Mobilenetv3 represents the backbone feature extraction network, Conv represents the convolutional layer, rate represents the dilation rate of dilated convolution, Image Pooling represents the image pooling operation, SE-Net represents the squeeze-excited attention module, Concat represents the stitching layer, Upsample represents the upsampling layer, and OUTPUT represents the output image.
[0063] Furthermore, the INPUT (input image) is first fed into Mobilenetv3 (backbone network) to complete the initial feature extraction. Its output features are divided into two paths: one path enters the parallel branch processing containing multi-dilation rate dilated convolution and image pooling, and the other path is directly passed. The features after branch processing are merged with the 1×1 Conv processing results, and then fed together with the features directly passed from Mobilenetv3 into SE-Net to optimize the channel weights. After adjusting the channels by 1×1 Conv, they are concatenated and fused with the initial features of Mobilenetv3 through Concat. The fused features are processed by Upsample×2, 1×1 Conv, and 3×3 Conv in sequence, and finally the resolution is restored by Upsample×4, and the OUTPUT (semantic segmentation result) is output.
[0064] S4: Construct the MS-DeepLabv3+ semantic segmentation network, which is an improvement on the original DeepLabv3+.
[0065] DeepLabv3+ is a semantic segmentation algorithm based on deep convolutional neural networks, consisting of an encoder and a decoder. The encoder employs depthwise separable convolution combined with dilated spatial pyramid pooling (ASPP). The former reduces computational complexity by splitting standard convolutions into depthwise convolutions and pointwise convolutions, while the latter utilizes dilated convolutions with different dilation rates to perform multi-scale sampling of feature maps, effectively capturing the contextual information of the target. The decoder, through bilinear interpolation upsampling, fuses the deep semantic features output by the encoder with shallow detail features, compensating for edge information lost during encoding and improving the accuracy of segmentation boundaries.
[0066] The MS-DeepLabv3+ semantic segmentation network is an improved multi-scale semantic segmentation model based on the classic DeepLabv3+ semantic segmentation network. It enhances the capture and fusion capabilities of multi-scale features by optimizing the encoder-decoder structure. The encoder, building upon the original DeepLabv3+'s depthwise separable convolution and dilated spatial pyramid pooling (ASPP) modules, introduces multi-scale grouped convolution or multi-branch feature extraction units. By setting parallel branches with different convolution kernel sizes or dilation rates, it achieves feature extraction of targets at different scales (such as crop seedlings and weeds of varying sizes) in the input image. The decoder employs lightweight upsampling and cross-layer feature concatenation strategies to accurately fuse the deep semantic features output by the encoder with shallow detail features. It can also embed attention modules such as squeeze excitation (SE) to adaptively enhance the response of key feature channels. It is widely applicable to high-precision semantic segmentation tasks such as UAV remote sensing image segmentation and medical image analysis.
[0067] It should be noted that, for the seedling region output by target detection, an MS-DeepLabv3+ semantic segmentation network based on DeepLabv3+ was constructed to achieve fine segmentation of the seedling contours. Semantic segmentation, as a pixel-level classification task, can assign each pixel in the image to a specific semantic category. In the grassland infrared thermal imaging scene, the temperature difference features between surviving seedlings and withered vegetation and soil, combined with the strong robustness of infrared images to changes in illumination, provide reliable support for accurate segmentation.
[0068] In one possible implementation, S4 specifically includes sub-steps S401 to S404: S401: Replace the Xception backbone network in the original DeepLabv3+ with MobileNetv3.
[0069] The Xception backbone network is a lightweight convolutional neural network architecture based on depthwise separable convolution. It decouples the traditional "channel convolution and spatial convolution" operations, performing spatial convolution on the feature map of each channel separately through channel-wise convolution, and then achieving information fusion between channels through pointwise convolution. This significantly reduces the number of model parameters and computational complexity while ensuring feature extraction capabilities. It employs a stacked structure of "linear bottleneck and residual connections," dividing the feature mapping process into multiple "inlet-intermediate-outlet" modules. The intermediate flow enhances feature extraction by repeatedly stacking the same depthwise separable convolution modules, while residual connections alleviate the gradient vanishing problem in deep networks, enabling the model to build deeper network layers.
[0070] MobileNetv3 integrates three key technologies: deep separable convolution, squeeze and activation (SE) lightweight attention module, and neural network architecture search (NAS). This significantly reduces the number of model parameters and inference time while maintaining detection and segmentation accuracy. The network architecture consists of a backbone feature extraction network and a neck feature fusion module. The backbone network uses deep separable convolution to decompose the channel and spatial operations of traditional convolution to reduce computation. SE modules are embedded in key feature layers to achieve channel-dimensional feature weighting. Simultaneously, NAS technology is used to search for the hardware-friendly optimal network structure. The activation function uses a lightweight h-swish function instead of the traditional ReLU, improving non-linear expressiveness while reducing computational overhead. At the network's end, a combination of average pooling and 1×1 convolution efficiently compresses the feature dimension, outputting highly discriminative deep semantic features.
[0071] S402: The dilated spatial pyramid pooling module in the original DeepLabv3+ uses dilated convolutions with different dilation rates.
[0072] The Spatial Pyramid Pooling (ASPP) module is a core module in semantic segmentation networks used to capture multi-scale contextual information. By setting dilated convolutions with different dilation rates, it performs multi-scale sampling of the input feature map without increasing the number of parameters or reducing the feature map resolution. It typically includes a 1×1 standard convolution and multiple parallel dilated convolution branches with different dilation rates (such as 6, 12, and 18), along with a global average pooling branch to obtain global contextual features. Finally, the output features of all branches are concatenated and fused, and then the channel dimension is adjusted by a 1×1 convolution to output a fused feature map containing local details, target features at different scales, and global semantic information.
[0073] Dilated convolution is an improved convolution operation that introduces holes (spaces) into the standard convolution kernel, expanding the receptive field of the convolution without increasing the number of parameters or computational complexity. The size of the space between kernel elements is controlled by setting the dilation rate. When the dilation rate is 1, dilated convolution is equivalent to standard convolution; when the dilation rate is greater than 1, the kernel inserts a corresponding number of zero-valued spaces between adjacent weight parameters, thus covering a larger area of the input feature map without changing the kernel size.
[0074] It should be noted that the Spatial Pyramid Pooling (ASPP) module is retained. By using dilated convolutions with different dilation rates to extract multi-scale seedling features, it can better capture the semantic information of seedlings of different sizes and improve the ability to perceive seedling boundaries and details.
[0075] Optionally, the ASPP module uses dilated convolutions with dilation rates of 6, 12, and 18.
[0076] S403: Insert a squeeze-excited attention module during the decoding stage of the original DeepLabv3+.
[0077] The Squeeze-Excitation (SE) module is a lightweight channel attention mechanism that enhances the network's feature representation capabilities by adaptively learning the importance weights of feature channels, strengthening the response of key feature channels, and suppressing interference from redundant channels. The module's operation consists of two steps: Squeeze and Excitation. First, global average pooling is used to "squeeze" the input feature map, compressing the two-dimensional features of each channel into a one-dimensional feature vector representing the global information of that channel. Then, an "excitation" mechanism is constructed using two fully connected layers to learn the weight coefficients of each channel. The first fully connected layer reduces the channel dimension to decrease computation, while the second fully connected layer restores the dimension and outputs weights between 0 and 1 using a sigmoid activation function. Finally, the weight coefficients are applied to the corresponding channels of the original feature map, completing the enhancement of key features.
[0078] It should be noted that a squeeze-excitement (SE) attention module is inserted in the decoding stage. By dynamically adjusting the channel weights of the feature map and setting the compression ratio to 16, the model is guided to focus on important features such as the seedling outline, thereby improving the accuracy and precision of semantic segmentation.
[0079] S404: Construct the MS-DeepLabv3+ semantic segmentation network based on MobileNetv3, dilated convolution, and squeezed-excitation attention module.
[0080] In this embodiment of the invention, a semantic segmentation network based on the original DeepLabv3+ and improved MS-DeepLabv3+ is constructed to achieve multi-dimensional performance optimization for fine segmentation of seedling contours in grassland infrared images. The encoder backbone network is replaced with the lightweight and efficient MobileNetv3, combined with depthwise separable convolutions, SE attention modules, and NAS architecture search technology. This significantly reduces the number of model parameters and inference time while ensuring the extraction capability of deep semantic features, adapting to the deployment requirements of edge devices such as drones. The ASPP module employs dilated convolutions with three different dilation rates (6, 12, and 18) to expand the receptive field without reducing feature map resolution, accurately capturing contextual information of seedlings at different scales and enhancing the perception of details and boundaries of seedling targets of varying sizes. In the decoding stage, a squeeze-excited attention module is inserted with a compression ratio of 16x. Through adaptive learning of feature channel weights, the response of key features such as seedling contours is dynamically enhanced, and redundant background interference such as soil and withered vegetation is suppressed, improving the fusion accuracy of deep semantic features and shallow detail features. The final MS-DeepLabv3+ network fully combines the strong robustness of infrared images to changes in illumination with the temperature difference characteristics between seedlings and the background to achieve pixel-level fine segmentation of the seedling area in the target detection output, providing high-precision technical support for grassland seedling survival status assessment and ecological monitoring.
[0081] S5: Input the labeled data from the standard grassland infrared image sample data into the CS-YOLOv7 target detection network and the MS-DeepLabv3+ semantic segmentation network respectively for training to obtain the seedling target detection model and the semantic segmentation model.
[0082] In this embodiment of the invention, standardized grassland infrared image annotation data is input into the CS-YOLOv7 target detection network and the MS-DeepLabv3+ semantic segmentation network for targeted training. This allows both networks to fully learn the characteristic patterns of seedling targets in the grassland infrared images. On one hand, the CS-YOLOv7 network accurately learns the mapping relationship between the location, bounding box, and survival status label of the seedling targets, and the trained seedling target detection model can quickly locate and identify seedlings in different states. On the other hand, the MS-DeepLabv3+ network masters the distinguishing features between the seedling outline and the background based on pixel-level annotation data, and the trained semantic segmentation model can achieve fine segmentation of the seedling region. Simultaneously, the standardized sample data ensures the stability and effectiveness of the training process, avoiding interference from noise and bias in the original data on model performance. The final detection and segmentation models possess high accuracy and strong robustness, and can collaboratively complete the task of accurate monitoring of grassland seedlings.
[0083] S6: Acquire infrared image data of the grassland to be detected.
[0084] S7: Input the infrared image data of the grassland to be detected into the seedling target detection model, and output the seedling location information and seedling survival status.
[0085] In this embodiment of the invention, by inputting the infrared image data of the grassland to be detected into the trained seedling target detection model, the model can quickly and accurately locate the position information of seedlings in the image by leveraging the multi-scale feature fusion and attention enhancement advantages of the CS-YOLOv7 network. Simultaneously, it efficiently determines the survival status of seedlings without requiring manual area-by-area inspection, significantly improving the efficiency and automation of grassland seedling monitoring. Furthermore, the model is trained based on standardized infrared image samples, exhibiting strong robustness to environmental temperature interference and image noise, and can adapt to the seedling monitoring needs of different grassland scenarios, providing accurate and reliable target area basic data for subsequent fine segmentation of seedling contours and grassland ecological assessment.
[0086] S8: The location information of the seedlings is segmented using a semantic segmentation model.
[0087] In this embodiment of the invention, the seedling location information output by the CS-YOLOv7 object detection model is input into the MS-DeepLabv3+ semantic segmentation model. Leveraging the model's multi-scale feature capture capabilities and channel attention mechanism, pixel-level fine segmentation is performed within the located seedling region. This accurately delineates the seedling contours and effectively distinguishes surviving seedlings from withered vegetation, soil, and other background elements. Furthermore, targeted segmentation based on object detection results avoids indiscriminate processing of the entire image, significantly reducing the model's computational overhead and improving segmentation efficiency. The final output seedling contour information, combined with the previously detected survival status labels, provides high-precision quantitative data support for assessing the growth status and coverage statistics of grassland seedlings.
[0088] S9: Calculate the survival rate of grassland seedlings based on the detection and segmentation results.
[0089] Optionally, the formula for calculating the survival rate of grassland seedlings is as follows: ; in, η Indicates the survival rate of seedlings in grassland. n Indicates the number of surviving grassland seedlings. m This indicates the total number of seedlings in the grassland.
[0090] In one possible implementation, after S9, the following is also included: The system overlays seedling location information, seedling survival status, grassland seedling survival rate, and infrared image data of the grassland to be tested to generate a visual evaluation report.
[0091] Furthermore, it also outputs detailed data tables (including area code, number of seedlings, survival rate, evaluation time, etc.).
[0092] Reference manual attached Figure 4 The diagram shows a structural schematic of a grassland restoration survival rate assessment system provided by the present invention.
[0093] The present invention also provides a grassland restoration survival rate assessment system 20, applied to the above-mentioned grassland restoration survival rate assessment method, comprising: Processor 201.
[0094] The memory 202 stores computer-readable instructions, which, when executed by the processor 201, implement the grassland restoration survival rate assessment method as described in the method embodiment.
[0095] The grassland restoration survival rate assessment system 20 provided by the present invention can perform the above-mentioned grassland restoration survival rate assessment method and achieve the same or similar technical effects. To avoid duplication, the present invention will not elaborate further.
[0096] It should be understood that the processor in the embodiments of the present invention can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0097] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0098] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0099] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0100] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0101] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0102] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0103] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0104] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0105] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0106] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0107] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0108] This invention provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the grassland restoration survival rate assessment method as described in the method embodiment.
[0109] The present invention provides a computer-readable storage medium that can implement the steps and effects of the grassland restoration survival rate assessment method in the above-described method embodiments. To avoid repetition, the present invention will not repeat them.
[0110] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0111] The following points need to be explained: (1) The accompanying drawings of the embodiments of the present invention only involve the structures involved in the embodiments of the present invention. Other structures can refer to the general design.
[0112] (2) For clarity, the thickness of layers or regions is enlarged or reduced in the drawings used to describe embodiments of the invention, i.e., these drawings are not drawn to scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element or there may be intermediate elements.
[0113] (3) Where there is no conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0114] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for assessing the survival rate of grassland restoration, characterized in that, include: S1: Collect sample data of infrared images of grassland; S2: Preprocess the grassland infrared image sample data to obtain standard grassland infrared image sample data; S3: Construct a CS-YOLOv7 object detection network based on the original YOLOv7; S4: Construct the MS-DeepLabv3+ semantic segmentation network, which is an improvement on the original DeepLabv3+; S5: Input the labeled data from the standard grassland infrared image sample data into the CS-YOLOv7 target detection network and the MS-DeepLabv3+ semantic segmentation network respectively for training to obtain the seedling target detection model and the semantic segmentation model; S6: Acquire infrared image data of the grassland to be detected; S7: Input the infrared image data of the grassland to be detected into the seedling target detection model, and output the seedling location information and seedling survival status; S8: The seedling location information is segmented using the semantic segmentation model; S9: Calculate the survival rate of grassland seedlings based on the detection and segmentation results.
2. The grassland restoration survival rate assessment method according to claim 1, characterized in that, The specific data of the grassland infrared image sample includes: temperature resolution, thermal sensitivity, image resolution, acquisition frame rate, emissivity, and object distance.
3. The grassland restoration survival rate assessment method according to claim 1, characterized in that, S2 specifically includes: S201: Perform temperature calibration on the image pixel values in the grassland infrared image sample data based on the ambient temperature parameters; S202: Denoising the temperature-calibrated grassland infrared image sample data by combining Gaussian filtering and median filtering; S203: Enhance image contrast of denoised grassland infrared image sample data through histogram equalization; S204: Label the grassland infrared image sample data after image contrast enhancement to obtain the standard grassland infrared image sample data.
4. The grassland restoration survival rate assessment method according to claim 3, characterized in that, The annotation content of the standard grassland infrared image sample data specifically includes: the bounding box of the seedling area, the outline of the seedling, and the survival status label.
5. The grassland restoration survival rate assessment method according to claim 1, characterized in that, S3 specifically includes: S301: Insert a convolutional block attention module after the C3 module of the original YOLOv7; S302: Insert a simple attention module after the fast spatial pyramid pooling module in the original YOLOv7; S303: Replace the detection head portion in the original YOLOv7 with bilinear interpolation; S304: Replace the LeakyReLU activation function in the original YOLOv7 with the Mish function; S305: Construct the CS-YOLOv7 object detection network based on the C3 module, the simple attention module, the bilinear interpolation, and the Mish function.
6. The grassland restoration survival rate assessment method according to claim 1, characterized in that, S4 specifically includes: S401: Replace the Xception backbone network in the original DeepLabv3+ with MobileNetv3; S402: The dilated spatial pyramid pooling module in the original DeepLabv3+ employs dilated convolution with different dilation rates; S403: Insert a squeeze-excited attention module during the decoding stage of the original DeepLabv3+; S404: Construct the MS-DeepLabv3+ semantic segmentation network based on the MobileNetv3, the dilated convolution, and the squeezed excitation attention module.
7. The grassland restoration survival rate assessment method according to claim 1, characterized in that, The specific formula for calculating the survival rate of grassland seedlings is as follows: ; in, η This indicates the survival rate of seedlings in the grassland. n Indicates the number of surviving grassland seedlings. m This indicates the total number of seedlings in the grassland.
8. The grassland restoration survival rate assessment method according to claim 1, characterized in that, Following S9, it also includes: The seedling location information, seedling survival status, grassland seedling survival rate, and infrared image data of the grassland to be tested are superimposed to generate a visual evaluation report.
9. A grassland restoration survival rate assessment system, characterized in that, include: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the grassland restoration survival rate assessment method as described in any one of claims 1 to 8.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the grassland restoration survival rate assessment method as described in any one of claims 1 to 8.