Deep learning-based method and system for forest leaf instance segmentation.

The LeafInst network enhances deep learning models for tree leaf instance segmentation by integrating adaptive spatial fusion and dynamic perception techniques, addressing data limitations and complexity, achieving superior accuracy and efficiency in forestry tasks.

JP7852966B1Active Publication Date: 2026-04-28NANJING FORESTRY UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NANJING FORESTRY UNIV
Filing Date
2025-12-12
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing deep learning models for tree leaf instance segmentation face challenges such as limited data availability, model generalization, slow processing speed, dependence on preprocessing, vegetation complexity, multiscale issues, morphological changes, lighting conditions, species diversity, and real-time scalability, which hinder accurate and efficient leaf segmentation in forestry applications.

Method used

A deep learning-based method using a LeafInst network with a progressive feature pyramid network, adaptive spatial fusion, dynamic asymmetric spatial perception, and top-down cascade decoder to enhance feature extraction and fusion, addressing multiscale and morphological changes, and a high-throughput leaf growth status index for quantitative evaluation.

Benefits of technology

The method significantly improves segmentation accuracy, computational efficiency, and generalization ability, enabling robust leaf instance segmentation across different species and lighting conditions, with a 7.1% improvement in segmentation mAP and reduced computational costs, supporting forestry applications like ecological monitoring and breeding selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007852966000001_ABST
    Figure 0007852966000001_ABST
Patent Text Reader

Abstract

Providing a deep learning-based method and system for forest leaf instance segmentation. [Solution] The method includes the steps of: acquiring vegetation images, inputting the vegetation images into a leaf instance segmentation model, and obtaining leaf instance segmentation prediction results; training the leaf instance segmentation model using a training set; generating dynamic fused features; acquiring a multi-source deformation feature layer corresponding to the dynamic fused features through a dynamic asymmetric spatial perception mechanism built into a dynamic anomalous regression head module; optimizing multi-scale features by employing a feature fusion strategy of a top-down cascade decoder module to obtain a multi-source fused feature layer; and further generating leaf instance segmentation prediction results using the multi-source fused feature layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing image processing technology, and particularly to a method and system for tree leaf instance segmentation based on deep learning.

Background Art

[0002] With the rapid development of remote sensing technology and artificial intelligence, tree leaf instance segmentation technology is playing an increasingly important role in fields such as ecological monitoring, plant phenotyping analysis, and forest resource surveys. However, existing leaf instance segmentation technologies still have the following drawbacks and problems.

[0003] Limit of data volume: Deep learning technology requires a large amount of data. In actual applications, it is still difficult to obtain large-scale and diverse remote sensing image data. Especially in annotation work that requires specialized knowledge such as leaves, the acquisition cost of annotation data is high, and the lack of sufficient training data affects the performance of deep learning models. Furthermore, due to the large morphological differences in leaves depending on tree species and growth seasons, more annotation data is required to cover these changes.

[0004] Limit of model generalization ability: Existing deep learning models are often trained on specific datasets, and their generalization ability for new remote sensing image data is limited, and there may be problems of overfitting. Leaves have large changes in morphology and texture features under different environmental conditions (different lighting, different weather, different seasons, etc.), so the adaptability of the model in actual applications is low. Therefore, in order to adapt to a wider range of remote sensing image application scenarios, it is necessary to further improve the generalization ability of the model.

[0005] Slow Processing Speed: Deep learning models typically require long training times and high computational resources, resulting in slow processing speeds in real-world applications. Forestry applications, in particular, require processing large volumes of drone remote sensing images, demanding high real-time performance. To improve the real-time performance and efficiency of the model, it is necessary to optimize the model structure and algorithms and reduce the computational complexity of the model.

[0006] Dependence on remote sensing image preprocessing: Since the quality and features of remote sensing images critically affect the density distribution estimation results, a series of preprocessing steps such as denoising, normalization, and enhancement must be performed on remote sensing images before training a deep learning model. These preprocessing steps introduce additional errors and uncertainties, affecting the accuracy of the model. Furthermore, these preprocessing steps increase the complexity and computational cost of the system.

[0007] Limitations regarding vegetation complexity: Conventional methods rely primarily on manually designed feature extractors, making it difficult to capture higher-dimensional features of vegetation density distributions. While deep learning techniques can automatically learn high-dimensional feature representations from data, they still have certain limitations and cannot handle the complexity of vegetation density distributions. Leaves possess complex morphological structures, texture features, and spatial distribution patterns, and this complexity presents significant challenges for instance segmentation.

[0008] Difficulty in addressing the multiscale problem in remote sensing images: Remote sensing images typically have features at multiple scales, such as detailed features and overall features, but deep learning models can usually only process information at a specific scale. Leaves have different scale levels, from single leaves to entire trees, and models need to be able to effectively process multiscale information. Therefore, when processing remote sensing images, it is necessary to consider how to effectively combine information from multiple scales and how to effectively transfer and merge information between different scales.

[0009] Challenges in leaf morphological changes: In natural environments, leaves are affected by various factors such as wind, gravity, and pests, causing deformation, bending, and overlapping. Due to these morphological changes, conventional detection methods based on geometric features have difficulty accurately identifying and segmenting leaf instances, requiring more robust feature extraction and segmentation algorithms.

[0010] Influence of lighting conditions: During the remote sensing image acquisition process, changes in lighting conditions (such as cloudy, sunny, or time of day) significantly affect leaf color, texture, and contrast features. Existing deep learning methods often show performance degradation when processing images under different lighting conditions, necessitating stronger lighting-invariant feature extraction capabilities.

[0011] Handling Species Diversity: In forestry, leaves of different tree species exhibit varying shapes, sizes, textures, and color characteristics, and this diversity presents challenges for general-purpose leaf instance segmentation models. Existing methods are often optimized for specific species or scenes, lacking the ability to handle leaves of different species universally.

[0012] Real-time and scalability requirements: Actual forestry applications often require processing large amounts of drone remote sensing images, placing high demands on the real-time and scalability of the system. Existing deep learning methods have problems such as low computational efficiency and high memory usage when processing large amounts of data, making it difficult to meet the requirements of actual applications. [Overview of the project] [Problems that the invention aims to solve]

[0013] To solve the problems of the conventional technology described above, the object of the present invention is to provide a forest leaf instance segmentation method and system based on deep learning. By performing leaf instance segmentation and downstream applications using an improved deep learning model, higher performance is provided in the fields of leaf segmentation and counting. [Means for solving the problem]

[0014] To achieve the above objectives, the present invention provides the following scheme.

[0015] A deep learning-based forest leaf instance segmentation method is: The process involves acquiring vegetation images, inputting the vegetation images into a leaf instance segmentation model, obtaining leaf instance segmentation prediction results, training the leaf instance segmentation model using a training set, and the training set includes a step that includes raw vegetation images. The process includes the steps of: extracting and enhancing features using the backbone module in the leaf instance segmentation model; integrating an adaptive spatial fusion mechanism in the incremental feature pyramid network to dynamically adjust feature weights and generate dynamic fused features; obtaining a multi-source deformed feature layer corresponding to the dynamic fused features through a dynamic asymmetric spatial perception mechanism built into the dynamic anomalous regression head module; optimizing multi-scale features by employing a feature fusion strategy of the top-down cascade decoder module to obtain a multi-source fused feature layer; and further generating the leaf instance segmentation prediction result using the multi-source fused feature layer.

[0016] To optionally generate the aforementioned dynamic fusion features, This method includes using a ResNet50 network pre-trained on an ImageNet model as a backbone network, integrating a progressive feature pyramid network, unifying the spatial resolution of different layers through an adaptive spatial fusion mechanism in the progressive feature pyramid network, dynamically adjusting the contribution of features from each layer to different layers by employing learnable weights, selecting leaf regions of interest, and obtaining the dynamically fused features.

[0017] Selectively acquiring the aforementioned dynamic fusion features includes the following equation:

number

[0018] Optionally, obtaining the multi-source deformation feature layer includes capturing the horizontal deformation feature of the dynamic fusion feature using the horizontal convolution kernel in the horizontal convolution branch, capturing the vertical deformation feature of the dynamic fusion feature through the vertical convolution kernel in the vertical convolution branch, capturing the overall contour feature of the dynamic fusion feature based on the standard convolution kernel in the depth convolution branch, and capturing the local detailed feature of the dynamic fusion feature by adopting the shallow convolution kernel in the shallow convolution branch; combining the horizontal deformation feature, the vertical deformation feature, the overall contour feature, and the local detailed feature along the channel dimension, and recovering the depth of the feature map using linear projection to obtain the multi-source deformation feature layer.

[0019] Selectively acquiring the multi-source fused feature layer is possible. The steps include: upsampling the high-level features in the multi-source deformation feature layer to the scale of the low-level features through transposed convolution, and performing dimensionality reduction processing using convolutional blocks; The process includes the steps of concatenating the high-level and low-level features after dimensionality reduction, recursively processing down to the lower layer, performing feature weighting using learnable convolution, and obtaining the multi-source fused feature layer.

[0020] Selectively obtaining the multi-source fused feature layer includes the following equation:

number

number

number

number

number

[0021] In the process of selectively acquiring the multi-source fused feature layer, the degree of feature reduction is controlled through multiple dynamic weight coefficients, and channel expansion is performed by employing depth-separable convolution, thereby aiming for downstream branches to equilibriumly select feature maps, and shallow features are reduced based on a dual-branch residual structure.

[0022] To optionally generate the leaf instance segmentation prediction results, The steps include: dynamically predicting the convolutional layer parameters of the MaskFCN unit based on the multi-source fused feature layer using the controller unit in the controller module; The method includes the steps of dynamically generating conditional instances using the aforementioned convolutional layer parameters, constructing a dynamic convolutional layer, performing instance prediction on the multi-source fused feature layer using the dynamic convolutional layer, and generating the leaf instance segmentation prediction results.

[0023] Optionally, the method further, The process includes establishing a leaf growth status index using an entropy weighting method combination of shape index and color index based on the leaf instance segmentation prediction results, performing a high-throughput automated analysis of the leaf phenotype using the leaf growth status index, and obtaining the analysis results. Here, the shape index includes length, width, perimeter, area, circularity, and rectangularity, and the color index includes the mean, median, and ternary values ​​of the RGB channels.

[0024] To achieve the above objectives, the present invention further provides a deep learning-based forest leaf instance segmentation system: An image acquisition subsystem for acquiring vegetation images, The system includes a leaf instance segmentation subsystem for inputting the vegetation images into a leaf instance segmentation model, obtaining leaf instance segmentation prediction results, and training the leaf instance segmentation model using a training set, wherein the training set includes original vegetation images. The backbone module in the leaf instance segmentation model is used to perform feature extraction and enhancement, and the adaptive spatial fusion mechanism in the incremental feature pyramid network is integrated to dynamically adjust feature weights and generate dynamic fused features. A multi-source deformed feature layer corresponding to the dynamic fused features is obtained through the dynamic asymmetric spatial perception mechanism built into the dynamic anomalous regression head module, and the feature fusion strategy of the top-down cascade decoder module is adopted to optimize multi-scale features and obtain a multi-source fused feature layer. Furthermore, the leaf instance segmentation prediction results are generated using the multi-source fused feature layer. [Effects of the Invention]

[0025] The beneficial effects of this invention are as follows:

[0026] This invention integrates a progressive feature pyramid network into the model neck, enables dynamic interaction of multiscale features through adaptive spatial fusion operations (ASFF), solves the problem of feature information loss in non-adjacent layers in conventional FPNs, dynamically adjusts the contribution of each feature layer at different layers through learnable weight parameters, and significantly improves the model's ability to sense changes in leaf scale.

[0027] To address the morphological deformation problem of forest leaves in natural environments caused by wind forces, this invention designs an innovative dynamic asymmetric spatial perception module. This module can effectively capture the irregular deformation features of leaves through four different types of convolutional operations: horizontal convolution, vertical convolution, depth convolution, and deformable convolution, overcoming the limitations of conventional convolutional operations in terms of local dependence.

[0028] This invention proposes a novel dynamic anomalous regression head architecture that employs a dual residual structure to recover features and effectively eliminate local field effects in convolutional operations. The controller unit dynamically predicts 169 convolutional layer parameters of the Mask FCN, enabling dynamic generation of conditional instances, ensuring accuracy while significantly reducing computational costs.

[0029] This invention re-examines the limitations of conventional feature fusion strategies and proposes an innovative TCFU strategy to replace conventional pixel-level additive fusion. This strategy assigns learnable convolution parameters to the feature weighting process through a recursive process, avoiding repeated weighting of low-level features, effectively solving the feature redundancy problem, and achieving hierarchical feature weight rebalancing.

[0030] Based on predicted instance segmentation results, the present invention proposes an innovative high-throughput leaf growth detailing index (LGCI) for quantitative evaluation of leaf phenotypic development. This index can effectively infer leaf growth trends, provide a scientific basis for breeding selection, and enable a shift from conventional qualitative analysis to quantitative and intelligent analysis.

[0031] In summary, this invention proposes a completely novel instance segmentation network, LeafInst, to address three challenges in leaf analysis in forest tree scenes: scale differences, lighting conditions, and species diversity. This network particularly improves the feature extraction capability for irregular leaf shapes: it enhances backbone feature recognition by adding a progressive pyramidal structure to the intermediate layer and employs dynamic asymmetric spatial perception techniques at the output end to effectively improve the accuracy of detailed recognition of multi-morphological leaves. Experiments have shown that this network is clearly superior to current mainstream algorithms in terms of segmentation effect. These techniques demonstrate good stability in practical applications. Among them, the LGCI index helps forestry workers to quickly and automatically evaluate the growth status of plantation forest leaves, provides data support for the selection and breeding of superior tree species, and promotes the development of smart forestry breeding. The specific technological advantages are as follows:

[0032] 1. Significant improvement in accuracy: On the Poplar-leaf dataset, LeafInst shows a 7.1% improvement in segmentation mAP compared to YOLOv11 and a 6.5% improvement compared to the transformer decoder MaskDino, demonstrating superior performance over large-scale visual models in zero-shot scenes.

[0033] 2. Optimization of computational efficiency: Through dynamic parameter prediction and TCFU strategies, we effectively solve the problems of redundancy of low-level features and iterative weighting in conventional methods, maintaining high accuracy while significantly improving computational efficiency.

[0034] 3. Strong generalization ability: The model demonstrates excellent performance in cross-dataset and cross-species transfer learning, and is directly applicable to other agricultural / forestry leaf analysis tasks, showing good generalization ability.

[0035] 4. High Practicality: The proposed LGCI index provides a quantifiable evaluation standard for forestry breeding, enabling a shift from conventional artificial evaluation to intelligent and standardized evaluation, and thus possesses significant practical application value. [Brief explanation of the drawing]

[0036] To more clearly illustrate embodiments of the present invention or technical solutions in the prior art, the drawings used in the embodiments are briefly introduced below. Clearly, the drawings in the following description represent only a few embodiments of the present invention, and those skilled in the art can obtain other drawings based on these without requiring any creative effort. [Figure 1] This is a flowchart of a deep learning-based forest leaf instance segmentation method according to an embodiment of the present invention. [Figure 2] This is a schematic diagram of a leaf instance segmentation model according to an embodiment of the present invention. [Figure 3] This is a workflow flowchart for a leaf instance segmentation model according to an embodiment of the present invention. [Modes for carrying out the invention]

[0037] The technical solutions in embodiments of the present invention will be clearly and completely described below with reference to the drawings of the embodiments. Clearly, the embodiments described are only some, and not all, embodiments of the present invention. All other embodiments that a person skilled in the art can obtain without creative work based on embodiments of the present invention are within the scope of the protection of the present invention.

[0038] To make the above-mentioned objectives, features, and advantages of the present invention clearer and easier to understand, the present invention will be described in more detail below with reference to the drawings and specific embodiments.

[0039] As shown in Figure 1, this embodiment discloses a forest leaf instance segmentation method based on deep learning: acquiring vegetation images, inputting the vegetation images into a leaf instance segmentation model, acquiring leaf instance segmentation prediction results, training the leaf instance segmentation model using a training set, the training set including raw vegetation images, and performing feature extraction and enhancement using the backbone module in the leaf instance segmentation model, integrating an adaptive spatial fusion mechanism in a progressive feature pyramid network to dynamically adjust feature weights and generate dynamic fused features, acquiring a multi-source deformed feature layer corresponding to the dynamic fused features through a dynamic asymmetric spatial perception mechanism built into the dynamic anomalous regression head module, optimizing multi-scale features by employing a feature fusion strategy of the top-down cascade decoder module to obtain a multi-source fused feature layer, and further using the multi-source fused feature layer to generate leaf instance segmentation prediction results.

[0040] Specifically, this embodiment discloses a deep learning-based forest leaf instance segmentation method, which includes: collecting raw vegetation image data to obtain a dataset, where the raw image data is in remote sensing image format; performing feature extraction and enhancement on the dataset, dynamically adjusting feature weights through a progressive feature pyramid network (AFPN) and adaptive spatial fusion operation (ASFF) to obtain dynamically fused features; obtaining a multi-source deformed feature layer using the built-in dynamic asymmetric spatial perception (DASP) technique based on a dynamic anomalous regression head (DARH) module; optimizing multiscale features and solving low-level feature redundancy problems to obtain a multi-source fused feature layer by employing a top-down cascaded decoder feature fusion (TCFU) strategy; and a controller unit predicting 169 convolutional layer parameters of Mask FCN, dynamically generating conditional instances, and performing final instance prediction using Mask FCN to obtain output results.

[0041] Furthermore, generating dynamically fused features involves using a ResNet50 network pre-trained on an ImageNet model as the backbone network, integrating a progressive feature pyramid network, unifying the spatial resolution of different layers through an adaptive spatial fusion mechanism within the progressive feature pyramid network, dynamically adjusting the contribution of features from each layer to different layers by employing learnable weights, selecting leaf regions of interest, and obtaining dynamically fused features.

[0042] Specifically, the gradual feature pyramid network (AFPN) is constructed as follows: ResNet50, pre-trained on ImageNet, is used as the backbone network, integrating the gradual feature pyramid network with an adaptive spatial fusion mechanism. Spatial resolutions of different layers are unified through an identity alignment method, and learnable weight parameters are employed to dynamically adjust the contribution of features from each layer to different layers. Leaf regions of interest are selected, and a multiscale feature pyramid is constructed.

[0043] Furthermore, obtaining a multi-source deformation feature layer includes: capturing the horizontal deformation features of the dynamic fusion feature using a horizontal convolution kernel during a horizontal convolution bifurcation; capturing the vertical deformation features of the dynamic fusion feature through a vertical convolution kernel during a vertical convolution bifurcation; capturing the overall contour features of the dynamic fusion feature based on a standard convolution kernel during a depth convolution bifurcation; and capturing the local detail features of the dynamic fusion feature by employing a shallow convolution kernel during a shallow convolution bifurcation; and concatenating the horizontal deformation features, vertical deformation features, overall contour features, and local detail features along the channel dimension, and recovering the depth of the feature map using linear projection to obtain a multi-source deformation feature layer.

[0044] Specifically, the Dynamic Asymmetric Spatial Perception (DASP) technique includes the steps of: extracting leaf morphological features in different directions through four different types of convolutional operations: horizontal convolution, vertical convolution, depth convolution, and deformable convolution; concatenating the output features of the four types of convolutional branches and recovering the depth of the feature map through linear projection; and designing a residual structure to preserve the original feature information and finally outputting a multi-source deformable feature layer.

[0045] Furthermore, obtaining a multi-source fused feature layer means that The method includes the steps of: upsampling the high-level features in a multi-source deformed feature layer to the scale of the low-level features through transposed convolution, and performing dimensionality reduction using convolutional blocks; and concatenating the high-level and low-level features after dimensionality reduction, recursively processing down to the lower layer, performing feature weighting using learnable convolution, and obtaining a multi-source fused feature layer.

[0046] Specifically, the top-down cascade decoder feature fusion (TCFU) strategy includes the steps of: upsampling high-level features to the scale of low-level features through transposed convolution; reducing feature redundancy through convolutional blocks; concatenating the dimensionality-reduced high-level and low-level features and processing recursively down to the lower layers; and weighting features through learnable convolutional parameters to avoid the problem of iterative weighting of low-level features.

[0047] Furthermore, in the process of acquiring the multi-source fused feature layer, the degree of feature reduction is controlled through multiple dynamic weight coefficients, and channel expansion is performed by employing depth-separable convolution. This aims to ensure that downstream branches equilibriumly select feature maps, and shallow features are reduced based on a dual-branch residual structure.

[0048] Specifically, the Dynamic Anomalous Regression Head (DARH) module further includes: controlling the degree of feature reduction through two dynamic weight coefficients, performing channel expansion using depth-separable convolution, ensuring that downstream branches equilibriumly select feature map portions, assisting the model in reducing shallow features through a dual branch residual structure, and including a dual residual structure to eliminate local field effects of convolution. “Multi-source fused feature layer”

number

[0049] Furthermore, the method further includes the steps of establishing a leaf growth status index using an entropy weighting method combination of shape index and color index based on the leaf instance segmentation prediction results, performing a high-throughput automated analysis of the leaf phenotype using the leaf growth status index, and obtaining the analysis results, where the shape index includes length, width, perimeter, area, circularity, and rectangularity, and the color index includes the mean, median, and ternary values ​​of the RGB channels.

[0050] Specifically, based on the prediction-based instance segmentation results, leaf growth status indices are established, and high-throughput automated analysis of leaf phenotypes is achieved through an entropy weighting method combination of shape indices and color indices. The shape indices include length, width, perimeter, area, circularity, and rectangularity, while the color indices include the mean, median, and ternary values ​​of the RGB channels.

[0051] This embodiment discloses a deep learning-based forest leaf instance segmentation method that, through an innovative network architecture design, can effectively address three challenges in forest leaf instance segmentation: scale changes, brightness changes, and shape changes. As shown in Figure 1, the overall method flow of the present invention includes five core steps: data acquisition (S1), feature extraction (S2), feature enhancement (S3), feature fusion (S4), and instance prediction (S5).

[0052] As shown in Figure 1, the method flow of the present invention begins with data collection, proceeds through feature extraction, feature enhancement, and feature fusion, and ultimately achieves instance prediction. Each stage has specific technical features and innovations, which together constitute a complete forest leaf instance segmentation system.

[0053] In the data collection stage (S1), the present invention uses a drone to collect raw image data of vegetation and obtain a high-quality dataset. Drone data collection has advantages such as low altitude, high resolution, scientific data support, and manual annotation optimization, enabling low altitude, high resolution, high speed, and precise sampling, providing scientific and accurate data support to this embodiment. Detailed manual annotation based on drone-sampling images ensures data accuracy and standardization, and establishes a standardized data quality control flow.

[0054] In the feature extraction stage (S2), the present invention uses a ResNet50FPN pre-trained on ImageNet as a skeletal model to select vegetation regions of interest through a multi-layer CA attention mechanism and construct a feature pyramid. The core technique in this stage is Adaptive Spatial Fusion Operation (ASFF), which first unifies feature maps of different hierarchical levels to the same resolution through an identity alignment method. Here, the downsampling operation employs a combination of max pooling and convolution, while the upsampling operation employs a bilinear interpolation method. ASFF dynamically adjusts the contribution of each feature layer to different hierarchical levels by learning a set of learnable weight parameters, and the weight parameters are normalized through a softmax function to achieve incremental multiscale feature fusion. By adopting an incremental fusion strategy, ASFF ensures that feature information is effectively transmitted between adjacent and non-adjacent hierarchical levels, solving the problem of feature information loss or degradation in conventional FPNs and significantly improving the model's ability to sense changes in leaf scale.

[0055] As shown in Figure 2, the algorithm model structure of the present invention represents the complete network architecture from the input image to the final segmentation result. Figure 2 shows the overall architecture of the algorithm of the present invention, which includes a feature extraction network, a feature enhancement module, a feature fusion strategy, and an instance prediction module. Each module is meticulously designed to ensure that it can effectively address various challenges in leaf instance segmentation.

[0056] The mathematical expression for the ASFF operation is as follows:

number

number

number

number

[0057] In the feature enhancement stage (S3), the present invention acquires a multi-source deformation feature layer through a Dynamic Asymmetric Spatial Perception (DASP) module. DASP is a branch specifically designed to capture feature extractions of different scales and shapes, and is designed for the morphological deformation problem of leaves in a natural environment caused by wind forces. This module includes four parallel feature extraction branches: a horizontal convolution branch captures the horizontal deformation features of the leaf using a horizontal convolution kernel; a vertical convolution branch captures the vertical deformation features of the leaf using a vertical convolution kernel; a depth convolution branch captures the overall contour features of the leaf using a standard convolution kernel; and a shallow convolution branch captures the local detail features of the leaf using a shallow convolution kernel. The convolution kernel sizes for horizontal and vertical convolution are affected by the resolution downsampling ratio (Stride), with models employing a smaller convolution kernel size K for higher downsampling ratios to ensure the accuracy of feature extraction. The output features of the four branches are fused through feature concatenation and linear mapping in the channel dimension, employing a residual structure to preserve the original feature information. Here, the residual connection strength is fixed at 0.3. The mathematical representation of the DASP module is as follows:

number

number

number

number

number

number

number

number

[0058] The mathematical representation of feature fusion is as follows:

number

[0059] The mathematical representation of the final residual structure is as follows:

number

number

number

number

[0060] In the feature fusion stage (S4), the present invention employs a top-down concatenated-decoder feature fusion (TCFU) strategy to optimize multiscale features. Conventional top-down pixel-level additive feature fusion has several problems, including: feature redundancy due to repeated weighting of low-layer features in the Neck stage and continued use of pixel-level additives in the Mask stage; the model ignoring overall contour features due to excessive weighting of low-layer texture details; and the lack of a learnable feature weight assignment mechanism. TCFU employs a recursive processing strategy, first performing dimensionality reduction using a C / 2 filter for each layer except the first layer, then upsampling the feature map using transposed convolution, and finally combining it with the features of the lower layer. TCFU assigns learnable convolution parameters to the feature weighting process, avoids the artificial a priori assumptions that result from simple additives, reduces feature redundancy through dimensionality reduction and upsampling, better balances the weight assignment of different hierarchical features, and achieves a rebalancing of hierarchical feature weights.

[0061] The mathematical representation of conventional pixel-level additive feature fusion is as follows:

number

number

number

number

[0062] In the instance prediction phase (S5), the present invention dynamically generates conditional instances through a controller unit and performs final predictions using MaskFCN. The controller unit is the core component of the instance prediction module and is specifically used to dynamically predict the convolutional layer parameters of MaskFCN and to enable the dynamic generation of conditional instances. The controller unit receives the output feature map of the DARH module and dynamically predicts the convolutional layer parameters of MaskFCN. A total of 169 parameters are predicted and divided into three sets of convolutional layer parameters to balance accuracy and computational cost. The controller unit dynamically generates conditional instances through the 169 predicted parameters, and the dynamic convolutional layer processes the feature map, achieving a transition from static to dynamic convolution. Each instance corresponds to a set of its own convolutional parameters, significantly reducing computational cost while maintaining high accuracy and supporting leaf instance segmentation of different shapes and scales.

[0063] As shown in Figure 3, the segmentation model of the present invention shows the complete flow from input features to the final segmentation mask. This model dynamically generates conditional instances through a controller unit and performs accurate instance segmentation prediction using MaskFCN.

[0064] Figure 3 shows in detail the core components of the segmentation model, namely key modules such as the controller unit, dynamic convolutional layers, and MaskFCN. Through the coordinated work of these modules, high-accuracy leaf instance segmentation is achieved.

[0065] The mathematical representation of the controller unit is as follows:

number

number

number

number

[0066] Centerness is used to measure the degree of deviation of each pixel from its target center. This helps the model sense the center of the corresponding leaf instance, thereby constraining candidate regions that are far from the target center. By complementing the box module with centerness, the cumbersome process of generating candidate anchor boxes of different scales for each pixel and filtering them by non-maximal suppression (NMS) is greatly simplified, and computational efficiency is improved. The mathematical representation of centerness is as follows:

number

[0067] This embodiment provides an application of a deep learning-based crop leaf instance segmentation method, demonstrating the broad application value of the present invention in the agricultural field. This application enables cross-domain application from leaves to crops through a zero-shot transfer learning method for crop leaf characteristics such as dense planting, leaf overlap, and lighting changes.

[0068] In agricultural applications, the main challenges faced by the present invention include the fact that crop leaves typically exhibit dense planting conditions, resulting in significant overlap and shading between leaves; that lighting conditions in agricultural environments change drastically, shifting from strong to weak light and from direct to scattered light; and that leaf morphological differences between different crop varieties are pronounced, ranging from narrow-leaved to broad-leaved crops, and from simple to compound leaves. To address these challenges, the present invention effectively captures leaf features of different shapes and scales through a Dynamic Asymmetric Spatial Perception (DASP) module, processes multiscale changes through Adaptive Spatial Fusion Operations (ASFF), and optimizes feature representation through a Top-Down Coupling-Decoder Feature Fusion (TCFU) strategy.

[0069] In the application validation of agricultural datasets, the present invention performs a zero-shot transfer learning test using the PhenoBench dataset. PhenoBench is a large, publicly available dataset specifically designed for semantic image interpretation in the agricultural field, providing RGB images recorded under real field conditions by drones equipped with high-resolution cameras. This dataset was obtained using a DJI M600 drone, acquiring motion-stabilized RGB images with a resolution of 11664 × 8750 pixels. After manual cropping and annotation by the team, each image patch size is 1024 × 1024 pixels, providing leaf segmentation instances of sugar beet crops and weeds.

[0070] Compared to forestry datasets, agricultural datasets have the following characteristics: agricultural datasets primarily provide orthorectified remote sensing images and have a single viewpoint, while forestry datasets capture leaf instances under natural lighting conditions and exhibit significant leaf overlap. Leaf distribution in agricultural environments is more dense, with a higher degree of occlusion and overlap between leaves. Leaf morphology of crops is relatively more regular, but with greater differences between varieties. The present invention's technology can be directly applied to the PhenoBench dataset without the need for specific adaptations for agricultural datasets through a zero-shot transfer learning method, fully demonstrating the versatility and robustness of the technology.

[0071] Technological innovations in agricultural applications include effectively processing different deformation features of agricultural leaves through the four-branch parallel architecture of the DASP module, achieving incremental multiscale feature fusion through ASFF technology to adapt to the diversity of leaf scales in agricultural environments, replacing conventional pixel-based addition with the TCFU strategy to reduce feature redundancy and improve segmentation accuracy in dense agricultural scenes, and dynamically generating conditional instances through the controller unit to support instance segmentation requests for different crop leaves.

[0072] The successful application of this embodiment demonstrates that the present invention is not only applicable to forest leaf instance segmentation, but also possesses strong cross-domain transfer capabilities, can effectively address various challenges in agricultural environments, and provides powerful technical support to fields such as precision agriculture, crop phenotypic analysis, and agricultural automation.

[0073] This embodiment verifies the effectiveness and superiority of the present invention's technical solution through detailed experiments. The experiments utilize the Poplar-leaf dataset, including leaf images under various environmental conditions, and evaluation metrics include segmentation mAP, detection accuracy, and computational efficiency. Compared to existing technologies, the present invention demonstrates superior performance in the following areas: In terms of improving segmentation mAP, it achieves a segmentation accuracy of 68.4%, an improvement of 7.1 points compared to YOLOv11's 61.3% and a improvement of 3.1 points compared to MaskDino's 65.3%. In terms of computational efficiency, it reduces feature redundancy through the TCFU strategy, improving processing speed. In terms of generalization ability, it demonstrates superior performance in cross-dataset and cross-species transfer learning. The mathematical representation of IoU (cross-ratio) used to evaluate algorithmic accuracy is as follows:

number

number

number

[0074] This embodiment further provides a forest leaf instance segmentation system based on deep learning, comprising an image acquisition subsystem for acquiring vegetation images, and a leaf instance segmentation subsystem for inputting the vegetation images into a leaf instance segmentation model, obtaining leaf instance segmentation prediction results, and training the leaf instance segmentation model using a training set. The training set includes raw vegetation images, and the system performs feature extraction and enhancement using the backbone module in the leaf instance segmentation model, dynamically adjusts feature weights by integrating an adaptive spatial fusion mechanism in a progressive feature pyramid network to generate dynamic fused features, obtains a multi-source deformed feature layer corresponding to the dynamic fused features through a dynamic asymmetric spatial perception mechanism built into the dynamic anomalous regression head module, optimizes multi-scale features by employing a feature fusion strategy in the top-down cascade decoder module to obtain a multi-source fused feature layer, and further generates leaf instance segmentation prediction results using the multi-source fused feature layer.

[0075] This embodiment further provides a deep learning-based forest leaf instance segmentation system: a collection module that collects raw vegetation image data to obtain a dataset, where the raw image data is in remote sensing image format; a model neck module that performs feature extraction and enhancement on the dataset and obtains dynamic fused features through a progressive feature pyramid network (AFPN) and adaptive spatial fusion operation (ASFF); a model head module that uses built-in dynamic asymmetric spatial perception (DASP) technology to obtain malformed leaf features through four different shapes of convolutions, design a dual branch residual structure to assist the model in reducing shallow features and ultimately obtain a multi-source deformation feature layer; a feature fusion module that uses a top-down cascade decoder feature fusion (TCFU) module to fuse and integrate information with the multi-source deformation feature layer to obtain a multi-source fused feature layer; and a controller unit that predicts 169 convolutional layer parameters of Mask FCN and dynamically generates conditional instances. This includes a prediction module that uses FCN to perform the final instance prediction and obtain the output result.

[0076] The phenotypic analysis module establishes a Leaf Growth Condition Index (LGCI) based on predicted instance segmentation results, enabling high-throughput automated analysis of leaf phenotypes. Through an entropy-weighted combination of shape and color indices, the phenotypic analysis module automatically identifies superior leaf morphologies, providing quantitative support for forestry breeding.

[0077] The system can achieve the following: adaptability to different lighting conditions, leaf morphology, and species type under zero-shot scenes; leaf instance segmentation at different scales such as single leaves, branched leaves, and whole trees; robust handling of nocturnal lighting insufficiency and leaf shape variations across different tree species.

[0078] The system is applicable to fields such as forestry resource surveys, plant phenotypic analysis, ecological monitoring, and precision agriculture, as well as to large-scale leaf instance segmentation and phenotypic parameter extraction through drone remote sensing images, and to providing automated technologies and quantitative support for forestry breeding and ecological research.

[0079] The embodiments described above are merely descriptions of preferred embodiments of the present invention and do not limit the scope of the invention. Any modifications and improvements made by those skilled in the art to the technical solutions of the present invention, without deviating from the spirit of the invention, should all be included within the scope of protection defined by the claims of the present invention.

Claims

1. A deep learning-based method for forest leaf instance segmentation, The process involves acquiring vegetation images, inputting the vegetation images into a leaf instance segmentation model, obtaining leaf instance segmentation prediction results, training the leaf instance segmentation model using a training set, and the training set includes a step that includes raw vegetation images. A deep learning-based forest leaf instance segmentation method comprising the steps of: extracting and enhancing features using the backbone module in the leaf instance segmentation model; integrating an adaptive spatial fusion mechanism in a progressive feature pyramid network to dynamically adjust feature weights and generate dynamic fused features; obtaining a multi-source deformed feature layer corresponding to the dynamic fused features through a dynamic asymmetric spatial perception mechanism built into the dynamic anomalous regression head module; optimizing multi-scale features by employing a feature fusion strategy of the top-down cascade decoder module to obtain a multi-source fused feature layer; and further generating the leaf instance segmentation prediction result using the multi-source fused feature layer.

2. Generating the aforementioned dynamic fusion features means A deep learning-based forest leaf instance segmentation method according to claim 1, characterized by using a ResNet50 network pre-trained with an ImageNet model as a backbone network, integrating a progressive feature pyramid network, unifying the spatial resolution of different layers through an adaptive spatial fusion mechanism in the progressive feature pyramid network, dynamically adjusting the contribution of features of each layer to different layers by employing learnable weights, selecting leaf regions of interest, and obtaining the dynamically fused features.

3. Obtaining the aforementioned dynamic fusion features involves the following equation: [Math 1] [Math 2] Here, [Math 3] and [Math 4] These are a set of learnable weight parameters, [Math 5] This is an identity alignment method, [Math 6] This is the i-th layer feature obtained through the AFPN neck, [Number 7] is X n From layer X i These are output features to the layer, [Number 8] is X i A deep learning-based forest leaf instance segmentation method according to claim 2, characterized in that the output features of the layer.

4. Obtaining the aforementioned multi-source deformation feature layer means that The steps include capturing the horizontal deformation features of the dynamic fusion feature using a horizontal convolution kernel during a horizontal convolution bifurcation, capturing the vertical deformation features of the dynamic fusion feature through a vertical convolution kernel during a vertical convolution bifurcation, capturing the overall contour features of the dynamic fusion feature based on a standard convolution kernel during a depth convolution bifurcation, and capturing the local detail features of the dynamic fusion feature by employing a shallow convolution kernel during a shallow convolution bifurcation. A deep learning-based forest leaf instance segmentation method according to claim 1, comprising the steps of: combining the horizontal deformation features, the vertical deformation features, the overall contour features, and the local detail features along the channel dimension, and recovering the depth of the feature map using linear projection to obtain the multi-source deformation feature layer.

5. A deep learning-based forest leaf instance segmentation method according to claim 1, characterized in that, in the process of acquiring the multi-source fused feature layer, the degree of feature reduction is controlled through a plurality of dynamic weight coefficients, channel expansion is performed by employing depth separable convolution, thereby aiming for downstream branches to equilibrium in selecting feature maps, and shallow features are reduced based on a dual branch residual structure.

6. The process of generating the aforementioned leaf instance segmentation prediction results is as follows: The steps include: dynamically predicting the convolutional layer parameters of the MaskFCN unit based on the multi-source fusion feature layer using the controller unit in the controller module; A deep learning-based forest leaf instance segmentation method according to claim 1, characterized by comprising the steps of: dynamically generating conditional instances using the aforementioned convolutional layer parameters, constructing a dynamic convolutional layer, performing instance prediction on the multi-source fused feature layer using the dynamic convolutional layer, and generating the leaf instance segmentation prediction result.

7. The method further, The process includes establishing a leaf growth status index using an entropy weighting method combination of shape index and color index based on the leaf instance segmentation prediction results, performing a high-throughput automated analysis of the leaf phenotype using the leaf growth status index, and obtaining the analysis results. A deep learning-based forest leaf instance segmentation method according to claim 1, characterized in that the shape index includes length, width, perimeter, area, circularity, and rectangularity, and the color index includes the mean, median, and ternary values ​​of the RGB channels.

8. A forest leaf instance segmentation system based on deep learning, An image acquisition subsystem for acquiring vegetation images, The system includes a leaf instance segmentation subsystem for inputting the vegetation images into a leaf instance segmentation model, obtaining leaf instance segmentation prediction results, and training the leaf instance segmentation model using a training set, wherein the training set includes original vegetation images. A deep learning-based forest leaf instance segmentation system according to any one of claims 1 to 7, characterized by performing feature extraction and enhancement using the backbone module in the leaf instance segmentation model, integrating an adaptive spatial fusion mechanism in a progressive feature pyramid network to dynamically adjust feature weights and generate dynamic fused features, obtaining a multi-source deformed feature layer corresponding to the dynamic fused features through a dynamic asymmetric spatial perception mechanism built into a dynamic anomalous regression head module, optimizing multi-scale features by employing a feature fusion strategy of a top-down cascade decoder module to obtain a multi-source fused feature layer, and further generating the leaf instance segmentation prediction result using the multi-source fused feature layer.

Citation Information

Patent Citations

  • Leaf vegetable image segmentation method, system, device and medium

    CN120612482A

  • Small sample learning crop canopy coverage calculation method and device based on background filtering

    JP2024117069A

  • Method and system for generating malnutrition maps

    JP2024527301A