GACL-DeepLabV3 +-based fine classification method for unmanned aerial vehicle remote sensing image citrus reiculata fruit tree plots

Through the GACL-DeepLabV3+ model, combined with lightweight networks and multi-scale feature fusion, the problems of blurred boundaries and feature overlap in the classification of Wogan fruit tree plots in drone remote sensing images were solved, achieving high-precision fruit tree plot identification and supporting the precise management of smart agriculture.

CN120635566AActive Publication Date: 2025-09-12GUILIN UNIVERSITY OF TECHNOLOGY

Patent Information

Application Number
CN202510752640.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-12
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Existing technologies have difficulty accurately distinguishing multi-age mixed-planting mandarin orange fruit tree plots in drone remote sensing images, especially when there are blurred boundaries, insufficient differentiation of features between plots of similar tree ages, overlapping features between young plots and bare soil areas, high model calculation complexity, and insufficient targeted data enhancement. These problems lead to insufficient classification accuracy and cannot meet the precise management needs of smart agriculture.

Method used

A method based on GACL-DeepLabV3+ is adopted, combined with the MobileNetV2 backbone network, the channel-aware lightweight attention module, the gated axial space attention module and the gated axial space pyramid structure. Through data enhancement and multi-scale feature fusion, lightweight and high-precision classification of Wogan fruit tree plots is achieved.

Benefits of technology

It has achieved lightweight deployment on the UAV platform, improved the average intersection-merger ratio of Wogan fruit tree plot classification, the classification accuracy of young plots and the pesticide patch recognition rate, significantly reduced the cost of manual surveys and improved orchard management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635566A_ABST
    Figure CN120635566A_ABST
Patent Text Reader

Abstract

The invention relates to an Orah fruit tree plot classification method in a multi-age mixed planting mode, and belongs to the technical field of agricultural remote sensing and image recognition. In order to solve the technical problem of difficulty in classification of citrus reiculata Blanco fruit tree plots in a multi-age mixed planting mode, remote sensing images of citrus reiculata Blanco planting areas are acquired through remote sensing of an unmanned aerial vehicle, and a multi-category data set including seedling stages, different tree ages and lime pesticide sprayed plots is constructed after preprocessing; a channel perception lightweight attention module (CLA) and a gating axial space attention module (GASA) are designed, a GACL-DeepLabV3 + model is constructed, the CLA is utilized to guide a key area to respond, spatial feature fusion is enhanced through the GASA, and fine classification of land parcels is realized. The method is mainly used for fine management of the citrus reiculata orchard, provides tree age distribution information for fruit farmers, assists in formulating differentiated management strategies such as water and fertilizer regulation and disease and pest early warning, and improves orchard monitoring efficiency and yield quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of agricultural remote sensing and image recognition, and specifically relates to a fine classification method for mandarin orange tree plots using UAV remote sensing images based on GACL-DeepLabV3+. Background Art

[0002] In the task of fruit tree plot classification in UAV remote sensing images, existing methods mainly focus on the overall segmentation of large-scale crops, but still face many challenges in the multi-age mixed planting scenario of mandarin orange trees. First, mandarin orange fruit tree plots of different ages are prone to blurred boundaries in remote sensing images. For example, the crowns of adjacent plots are staggered or the branches and leaves overlap, making it difficult for traditional segmentation algorithms to accurately define the plot outline. This problem stems from the complexity of the fruit tree canopy structure and the limitations of the resolution of UAV images. Especially in low-altitude images, the detailed information of densely planted areas is easily disturbed by factors such as lighting and shadows, further exacerbating the error of boundary segmentation. Secondly, mandarin orange trees at different growth stages (such as 1-year-old and 2-3-year-old plots) show similar spectral and texture characteristics in the images, and conventional models find it difficult to capture their subtle differences. This type of confusion is closely related to the gradual changes in canopy morphology during the fruit tree growth cycle, especially the small differences in canopy height and branch density between adjacent tree ages, resulting in insufficient feature discrimination. Furthermore, plots of young fruit trees (e.g., in the seedling stage) have low vegetation cover, and the exposed surface and surrounding non-planted areas (e.g., dirt roads and wasteland) are prone to spectral overlap in imagery, making it difficult for existing methods to effectively distinguish such plots. The root cause of this problem is that the vegetation signal in seedling-stage plots is weak, while surface background noise dominates. Traditional models lack the ability to collaboratively model local and global features, making it difficult to suppress background interference.

[0003] Existing technologies face the following difficulties in addressing these issues: First, traditional convolutional neural networks rely on fixed receptive fields to extract features, making it difficult to adaptively adjust to multi-scale contextual information. For example, atrous spatial pyramid pooling (ASPP) captures multi-scale features through atrous convolutions with a preset dilation rate. However, its fixed structure is not adaptable to the morphological diversity of mandarin orange orchards, and is particularly prone to introducing redundant information in boundary regions. Second, existing attention mechanisms struggle with lightweightness and computational efficiency. For example, while global attention can enhance feature relevance, its high computational complexity makes it difficult to deploy in real-time monitoring systems on drone platforms. Furthermore, lightweight designs often result in the loss of critical spatial information, impacting classification accuracy. Third, data augmentation strategies lack specificity. Existing methods typically employ general augmentation techniques (such as random flipping and cropping), but lack targeted simulation of noise types unique to mandarin orange orchards (such as local spectral anomalies caused by pesticide spraying), resulting in insufficient robustness in complex scenarios. Fourth, the uneven distribution of samples across plots of different tree ages in mixed planting patterns is a prominent issue. For example, the proportion of plots with old trees is high while the samples of plots with young trees are scarce. Traditional data balancing methods (such as oversampling) are prone to introduce overfitting risks, further exacerbating classification errors.

[0004] These challenges make it difficult for existing technologies to meet the precise management needs of multi-age mandarin orange orchards in smart agriculture, such as differentiated water and fertilizer regulation and pest and disease monitoring. Addressing these challenges requires balancing model lightweighting, multi-scale feature fusion, and noise robustness. Furthermore, efficient data augmentation and sample balancing strategies must be designed, placing high demands on innovative algorithmic architectures and optimized computing resources. Summary of the Invention

[0005] An object of the present invention is to solve at least the above problems and to provide at least the advantages which will be described hereinafter.

[0006] The present invention also solves the following technical problems:

[0007] Existing technologies lack sufficient accuracy for classifying plots of mixed-age mandarin orange trees. Traditional remote sensing image processing methods struggle to effectively distinguish between plots of varying tree ages and conditions, and the large number of model parameters makes real-time deployment on drone platforms impossible. Mixed planting patterns lead to blurred plot boundaries and overlapping features between younger plots and bare soil areas, necessitating a lightweight, high-precision classification method.

[0008] Traditional fixed-size image cutting can easily destroy the canopy continuity of Wogan fruit tree plots (such as branch breakage in the seedling stage and pesticide patch segmentation), reduce the quality of labeled data, and cause the model to be insufficiently sensitive to complex boundaries.

[0009] The traditional attention mechanism has high computational complexity and lacks joint modeling of the correlation between channel semantics and spatial position, resulting in high-level feature redundancy and difficulty in distinguishing subtle differences between plots of trees with similar ages.

[0010] The canopy structure and branch density of mandarin orange trees of different ages under mixed planting patterns are significantly different. Traditional convolution is difficult to adaptively capture axial spatial features, and boundary segmentation is easily affected by background interference.

[0011] There are significant scale differences among multi-age mandarin orange plots (fine texture in the seedling stage vs. wide distribution in the adult stage). Traditional hollow pyramids rely on a fixed expansion rate and cannot dynamically adapt to the gradual changes in tree age, resulting in the failure of multi-scale feature fusion.

[0012] Shallow detail features (branch density) and deep semantic features (canopy layout) are difficult to align due to resolution differences. Cross-scale feature fusion is insufficient in mixed planting scenarios, resulting in a high boundary missegmentation rate.

[0013] The morphological complexity of multi-age mixed-crop plots makes it impossible for a single feature extraction network to take into account both local details and global layout. The model easily confuses young plots with bare soil areas, and the classification results lack application guidance value.

[0014] To achieve the above-mentioned object of the present invention, the present invention provides a method for fine classification of Wogan fruit tree plots based on UAV remote sensing images using GACL-DeepLabV3+, comprising the following steps:

[0015] S1: Use UAV to obtain visible light remote sensing images of mandarin orange fruit trees of different ages and conditions in the study area;

[0016] S2: Preprocessing the acquired remote sensing images to generate centimeter-level orthophotos. The preprocessing includes feature extraction and aerial triangulation to generate dense point clouds, digital surface model construction, and orthorectification.

[0017] S3: Construct a dataset of mandarin orange tree plots based on orthophotos. The dataset includes plots in the seedling stage, plots with 1-year-old fruit trees, plots with 2-3-year-old fruit trees, plots with 4-5-year-old fruit trees, and plots with lime-sprayed fruit trees. The images are annotated using manual annotation tools and then segmented.

[0018] S4: Perform data augmentation on the dataset, including applying Gaussian noise with an intensity of 10-50 and rotation transformation to the image to generate a data augmented training set with noise robustness and multi-angle features, and divide it into training set, validation set and test set in proportion;

[0019] S5: Construct a GACL-DeepLabV3+ model, which includes a MobileNetV2 backbone network, a channel-aware lightweight attention module, a gated axial spatial attention module, and a gated axial spatial pyramid structure. Train the GACL-DeepLabV3+ model using the data augmentation training set. During training, a stochastic gradient descent optimizer is used with an initial learning rate of 7e-3, a cosine annealing learning rate decay method, a batch size of 4, and 100 training epochs.

[0020] S6: Input the test set into the trained GACL-DeepLabV3+ model and output the classification results of the Wogan fruit tree plots.

[0021] Preferably, the operation of the drone of the present invention to obtain visible light remote sensing images further includes: establishing a flight parameter configuration table based on a preset fruit tree age gradient database and canopy morphological feature library, using a multi-rotor drone equipped with a visible light multispectral imager with a polarization filtering function, and performing aerial photography with the canopy as the reference plane during the morning scattered light-dominated period, wherein young trees with a tree age of less than 3 years are flown at a flight altitude of 250 meters and a heading overlap rate of 80%, and adult and high-yield trees are flown at a flight altitude of 250 meters and a heading overlap rate of 70%, and an edge computing unit-based real-time image quality diagnosis module is configured to perform online detection and re-photography decisions on leaf texture clarity and abnormal fruit coloring areas.

[0022] Preferably, the manual labeling tool of the present invention is Labelme, and the labeled image segmentation adopts an adaptive segmentation algorithm based on saliency guidance.

[0023] Preferably, the channel-aware lightweight attention module of the present invention extracts the channel semantic information of the input features through global average pooling, generates a channel semantics-guided Key vector, and generates a spatial Query matrix through linear transformation. The Key vector and the Query matrix are dot-product normalized to generate a spatial attention map, which is then weighted and output through residual connection.

[0024] Preferably, the gated axial spatial attention module of the present invention enhances the spatial feature modeling capability of the Wogan fruit tree plots through axial convolution and gating mechanisms, specifically including: dividing the input features into two parts along the channel dimension, applying depth-separable convolution along the longitudinal and transverse axes to the first C channel features, respectively, to capture the canopy structure and branch density differences of the Wogan fruit tree plots in the vertical and horizontal directions, generating axial response features and summing them, and adaptively modulating the response features of the last C channel features through learnable gating parameters to suppress background interference and enhance the boundary features of plots of different tree ages, and then outputting the fused features through residual connections, thereby solving the problem of blurred boundaries of Wogan fruit tree plots due to mixed planting and overlapping features of young plots and non-plot areas.

[0025] Preferably, the gated axial space pyramid structure of the present invention constructs a multi-scale spatial modeling mechanism for Wogan fruit tree plots by layering the gated axial space attention modules, specifically including: processing the input features through multiple parallel branches, each branch compressing the channel dimension through 1×1 convolution and then introducing 1-4 layers of the gated axial space attention modules, and gradually expanding the receptive field using modules of different layers, wherein the shallow branches focus on the fine-grained canopy texture of the Wogan seedling plots through single-layer modules, and the deep branches capture the wide-area canopy distribution characteristics of the 4-5 year old plots through multi-layer modules; the multi-scale features output by the module are unified in resolution through bilinear interpolation, and then spliced ​​with the shallow detail features extracted by the MobileNetV2 backbone network, fusing the local branch density information and the global spatial layout characteristics of the Wogan fruit tree plots, thereby solving the boundary fuzzy problem caused by morphological differences of plots of different ages under the multi-age mixed planting mode, and suppressing the misclassification interference of young plots and bare soil areas.

[0026] Preferably, the fusion mechanism of the multi-scale features and shallow features output by the gated axial spatial pyramid structure of the present invention is designed for the morphological complexity of the Wogan fruit tree plot, specifically comprising the following steps:

[0027] Step 1: Shallow feature extraction and retention: Based on the shallow features extracted by the MobileNetV2 backbone network, the fine-grained branch density distribution characteristics and canopy edge texture details of the Wogan seedling plots are retained to enhance the ability to distinguish between young fruit tree plots and bare soil and road non-plot areas;

[0028] Step 2: Multi-scale feature alignment and adaptation: A bilinear interpolation algorithm is used to upsample the multi-scale features output by the gated axial space pyramid (shallow branches: fine textures in seedlings; deep branches: wide-area canopy layers in 4-5 year old trees) to make their spatial resolution consistent with the shallow features extracted by the MobileNetV2 backbone network (e.g., 128×128), thereby achieving geometric alignment of cross-scale features.

[0029] Step 3: Cross-scale feature fusion and boundary reinforcement: The aligned multi-scale features and shallow features are spliced ​​along the channel dimension, and cross-scale information fusion is performed through two 3×3 convolutional layers:

[0030] First convolution: Extract the branch density gradient variation characteristics within the Wogan fruit tree plots and enhance the differential response of canopy cover between plots of different tree ages;

[0031] Secondary convolution: Captures the patchy distribution characteristics of plots sprayed with lime and pesticides and the scattered boundary mutation signals of young fruit tree plots, and suppresses the missegmentation of adjacent plots in mixed planting patterns due to canopy overlap or background interference;

[0032] Step 4: Reconstructing a high-resolution classification map: The fused features are restored to the original image resolution through four-fold upsampling to generate a classification result map that matches the pixel-level of the drone orthophoto. This accurately marks the continuous boundaries and internal structure of the mandarin orange tree plots of different ages, solving the problems of blurred boundaries caused by mixed planting of multiple ages and missed detection of transition areas between young plots and bare soil.

[0033] Preferably, the construction of the GACL-DeepLabV3+ model of the present invention is designed for a multi-age mixed planting pattern of Wogan fruit tree plots, specifically comprising the following steps:

[0034] Step 1: Multi-level feature extraction of input images: The visible light image of the Wogan fruit tree plot acquired by the drone is input into the MobileNetV2 backbone network to extract shallow features and high-level features respectively:

[0035] Shallow layer characteristics: retain the fine-grained shoot density distribution, canopy edge texture of seedling-stage plots, and soil characteristics of exposed surfaces of young plots to distinguish non-fruit tree areas;

[0036] High-level features: Capture the wide-area canopy distribution pattern of 4-5 year old plots and the patchy spectral response of pesticide-sprayed plots, characterizing the global spatial layout of mature Wogan fruit trees;

[0037] Step 2: Channel-aware attention guides spatial feature enhancement: The channel-aware lightweight attention module is embedded after the high-level features to optimize the semantic distinction ability of the Wogan plot classification through the following operations:

[0038] Extract channel semantic information of high-level features based on global average pooling and generate channel semantic guided key vectors;

[0039] Convert the spatial location information into a query matrix and calculate the correlation weight between channel semantics and spatial location;

[0040] Residual connections are used to enhance the response intensity to sparse shoot areas in seedling plots and local patches of pesticide spraying plots, thereby suppressing background noise interference.

[0041] Step 3: Multi-scale spatial modeling and tree age feature adaptation: The high-level features enhanced by CLA are input into the gated axial spatial pyramid structure to adapt to the age differences of the Wogan fruit tree plots through the following mechanism:

[0042] Shallow branching focused on the fine-grained canopy texture of the seedling-stage plots, enhancing the contrast between young fruit trees and the bare soil boundary;

[0043] Deep branches expand the receptive field, capturing the wide-area canopy distribution of 4-5 year old tree plots and the staggered boundary characteristics between adjacent plots;

[0044] The output multi-scale features are aligned by bilinear interpolation to preserve the morphological gradient characteristics of plots of different tree ages;

[0045] Step 4: Classification result generation and application optimization: The cross-scale fused features are reconstructed through a high-resolution method to generate a classification result map that matches the pixel level of the drone orthophoto. The continuous boundaries and transition areas of the multi-age mandarin orange tree plots are accurately marked, and the results are imported into the orchard intelligent management system to guide differentiated fertilization, pest and disease control, and harvest planning, significantly reducing manual survey costs and improving orchard management efficiency.

[0046] The present invention has at least the following beneficial effects:

[0047] Through end-to-end lightweight model design (MobileNetV2+attention module), an average intersection over union (mIoU) of 94.03% is achieved with only 5.8MB of parameters, significantly reducing the cost of drone deployment. Combining multi-scale feature fusion and noise-robust data enhancement, the problem of blurred boundaries in mixed planting patterns is solved, the classification accuracy of young plots is improved by 10%, and the pesticide patch recognition rate is increased to 95%, providing reliable technical support for precise orchard management.

[0048] The CLA module guides spatial attention through channel semantics, with a significantly lower parameter count than traditional attention mechanisms, improving computational efficiency. In the task of classifying plots of trees with similar ages (2-3 years vs. 4-5 years), it improves the discrimination index and effectively reduces feature redundancy.

[0049] The GASA module captures canopy density gradients through axial convolution, improving the intersection-over-union (IoU) of boundary segmentation; the gating mechanism suppresses background interference and reduces the misclassification rate of young plots and bare soil areas, significantly enhancing the robustness of the model in mixed planting scenarios.

[0050] The gated axial pyramid structure of the layered stacked GASA modules adapts to the morphological gradient of plots from seedling to adult stage, and the average recall rate (Recall) in multi-age mixed scenarios reaches 96.7%; the multi-layer branch structure expands the receptive field, deeply extracts data texture features, improves the classification accuracy of the seedling stage, and reduces the segmentation error of the adult stage.

[0051] The cross-scale feature fusion mechanism aligns the resolution through bilinear interpolation to reduce the boundary missegmentation rate; the secondary convolution layer strengthens the patch response of pesticide spraying, and four times upsampling achieves pixel-level matching to reduce the missed detection rate in the transition area.

[0052] The fusion of multi-level feature extraction and attention guidance enables the model to achieve an overall accuracy (OA) of 96.7% in complex orchard scenarios; after the classification results are imported into the management system, the efficiency of differentiated fertilization is improved, the cost of manual survey is reduced, and the efficiency of orchard management is significantly improved.

[0053] Other advantages, objectives and features of the present invention will be reflected in part from the following description and will be understood by those skilled in the art through study and practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 Schematic diagram of the study area of ​​the present invention, (a) is the location indication map of the study area, (b) is the drone image of the study area, and (c) is the detailed map of the Wogan fruit tree plot;

[0055] Figure 2 This is a classification diagram of the Wogan fruit tree plots of the present invention, (a) is a seedling plot, (b) is a 1-year-old plot, (c) is a 2-3-year-old plot, (d) is a 4-5-year-old plot, and (e) is a lime-sprayed agricultural plot;

[0056] Figure 3 This is a data enhancement example of the present invention;

[0057] Figure 4 This is the structure diagram of the channel-aware lightweight attention (CLA) of the present invention;

[0058] Figure 5 This is a structural diagram of the gated axial spatial attention module (GASA) of the present invention;

[0059] Figure 6 This is a structural diagram of the Gated Axial Space Pyramid (GASP) of the present invention;

[0060] Figure 7 This is the GACL-DeepLabV3+ model structure diagram of the present invention;

[0061] Figure 8The classification and recognition results of seedling plots in the present invention are shown in Figure 1. (a) is the original image, and (b) is the result image.

[0062] Figure 9 The classification and recognition results of the one-year-old tree plots of the present invention are shown in Figure 1. (a) is the original image, and (b) is the result image.

[0063] Figure 10 The classification and recognition results of the 2-3 year old tree plots in this invention, (a) is the original image, (b) is the result image;

[0064] Figure 11 The classification and recognition results of the 4-5 year old tree plots in this invention, (a) is the original image, (b) is the result image;

[0065] Figure 12 This is the classification and recognition result of the plot sprayed with lime pesticide according to the present invention, (a) is the original image, and (b) is the result image. DETAILED DESCRIPTION

[0066] The present invention will be described in further detail below in conjunction with the accompanying drawings so that those skilled in the art can implement the invention with reference to the description.

[0067] It should be understood that terms such as “having”, “including” and “comprising” used herein do not preclude the existence or addition of one or more other elements or combinations thereof.

[0068] It should be noted that the experimental methods described in the following embodiments are conventional methods unless otherwise specified, and the reagents and materials can be obtained from commercial channels unless otherwise specified.

[0069] <Example>

[0070] A fine classification method for mandarin orange tree plots based on UAV remote sensing images using GACL-DeepLabV3+ is proposed, which includes the following steps:

[0071] S1: Use UAV to obtain visible light remote sensing images of mandarin orange fruit trees of different ages and conditions in the study area;

[0072] S2: Preprocessing the acquired remote sensing images to generate orthophotos with a resolution of 0.05-0.10 meters. The preprocessing includes feature extraction and aerial triangulation to generate dense point clouds, construction of digital surface models, and orthorectification.

[0073] S3: Construct a dataset of mandarin orange tree plots based on orthophotos. The dataset includes plots in the seedling stage, plots with 1-year-old fruit trees, plots with 2-3-year-old fruit trees, plots with 4-5-year-old fruit trees, and plots with lime-sprayed fruit trees. The images are annotated using manual annotation tools and then segmented.

[0074] S4: Perform data augmentation on the dataset, including applying Gaussian noise with an intensity of 10-50 and rotation transformation to the image to generate a data augmented training set with noise robustness and multi-angle features, and divide it into training set, validation set and test set in proportion;

[0075] S5: Construct a GACL-DeepLabV3+ model, which includes a MobileNetV2 backbone network, a channel-aware lightweight attention module, a gated axial spatial attention module, and a gated axial spatial pyramid structure. Train the GACL-DeepLabV3+ model using the data augmentation training set. During training, a stochastic gradient descent optimizer is used with an initial learning rate of 7e-1, a cosine annealing learning rate decay method, a batch size of 4, and 100 training epochs.

[0076] S6: Input the test set into the trained GACL-DeepLabV3+ model and output the classification results of the Wogan fruit tree plots.

[0077] The drone can be a multi-rotor model, with an imager that supports multispectral imaging. The edge computing unit can be integrated into the drone's control system. The drone's fuselage can be made of lightweight carbon fiber, and its sensors can use silicon-based CMOS chips. The imager is mounted on the drone's gimbal, and the edge computing unit is connected to the flight control module via a CAN bus.

[0078] Workflow: The drone automatically plans its flight path based on preset parameters and transmits images to the edge computing unit in real time during flight. The edge computing unit assesses image quality and immediately generates re-photographing instructions if blur or missing target categories are detected. Preprocessing software performs aerial triangulation through feature matching and bundle adjustment, ultimately outputting orthophotos with a resolution of 0.068 meters.

[0079] The image size was set to 512×512 pixels. Data augmentation included applying Gaussian noise with intensities of 10, 25, and 40, as well as 90°, 180°, and 270° rotations. The final dataset was split into training, validation, and test sets with a ratio of 8:1:1.

[0080] The labeling tool can be Labelme open-source software, running on Anaconda Prompt. Data can be stored on SSDs, and the saliency detection model can be deployed in the PyTorch framework. The segmentation model is also implemented in PyTorch.

[0081] Workflow: The data augmentation module applies noise and rotation to the image to generate an augmented dataset containing 2032 slices, which is stored on the local server.

[0082] The GACL-DeepLabV3+ model uses MobileNetV2 as the backbone network. The Channel-Aware Attention Module (CLA) extracts channel semantic information through global average pooling, generates a key vector and a spatial query matrix, and then normalizes the dot product to form an attention map. The Gated Axial Spatial Pyramid (GASP) consists of four branches, each stacking 1-4 layers of Gated Axial Attention Modules (GASA). Training uses a stochastic gradient descent optimizer with an initial learning rate of 7e-3, a batch size of 4, and 100 epochs. The loss function is cross-entropy.

[0083] Equipment and Materials: Model training is based on the PyTorch framework and runs on a server equipped with NVIDIA A100 GPU.

[0084] Workflow: The augmented training set is fed into the model, forward propagation is used to calculate the loss, and backward propagation is used to update the parameters. During training, accuracy is evaluated on the validation set every 10 epochs, and the optimal model weights are saved. During testing, 512×512 slices are fed into the model, and the model outputs a classification result image. This image is then upsampled fourfold to restore it to its original resolution, and the classification result is finally output.

[0085] Technical effects:

[0086] Dataset construction and data enhancement: saliency-guided cutting preserves key features, buffer repair reduces edge distortion, and noise and rotation enhancement improves model robustness.

[0087] Model construction and training: Lightweight design reduces computational overhead, and multi-scale attention mechanism enhances boundary segmentation capabilities.

[0088] In another technical solution, in order to take into account the computational efficiency of the model and explore the synergy between channel and spatial features, the present invention proposes a channel-aware lightweight attention module (CLA). CLA uses the global channel semantics of the input features as guiding information, and generates spatial attention weights with channel perception capabilities by calculating the correlation between channel semantics and spatial position, thereby achieving response enhancement in spatially salient areas.

[0089] The structure of CLA is as follows Figure 4 As shown, the feature map x∈R C×H×W First, global average pooling is performed along the spatial dimension to obtain the overall response of each channel, which is expressed as:

[0090] This vector is transformed by the linear transformation w k ∈R C'×C Get the channel semantics guided Key, expressed as K∈W k ·K c ∈RC′ At the same time, the input feature x is flattened to x * ∈R (H·W)×C , and through the linear transformation w q ∈R C'×C Get the Query matrix Q∈W q ·x * R C′×(HvW) . Do the dot product of the channel information Key and the spatial Query vector and normalize them to get the spatial attention map:

[0091] A=soft max(K T ·Q)∈R 1×(H·W) (1)

[0092] Finally, the original input feature map is weighted by element multiplication and the final result is output through the residual connection:

[0093] y=γ·A☉x+x (2)

[0094] where γ∈R is a learnable scaling factor.

[0095] CLA introduces a simple Query-Key process and uses channel-aware information to guide spatial features, thereby strengthening the model's responsiveness to important spaces and achieving effective semantic enhancement while reducing parameters.

[0096] Workflow: After the attention map is element-wise multiplied with the original feature map, the weighted intensity is adjusted by the scaling factor and then added to the original feature map. The output feature map retains the original information while enhancing the response of the salient areas.

[0097] Technical effects:

[0098] Channel-level semantic information is extracted through global average pooling, enhancing the model's sensitivity to key channels.

[0099] The query-key interaction mechanism is used to dynamically allocate spatial weights and reduce redundant feature interference.

[0100] A learnable scaling factor is used to balance the original features and the attention-weighted results to avoid gradient disappearance in the early stages of training.

[0101] In another technical solution, to enhance the model's ability to model spatial structure and target category boundaries in semantic segmentation tasks, the present invention proposes a lightweight gated axial spatial attention (GASA). The spatial attention mechanism strengthens the characteristics of the target area and improves the ability to capture target details by focusing on key spatial information. The core of GASA is to combine the idea of ​​axial attention on the basis of spatial attention and introduce learnable gating parameters to enhance the perception of local and long-range spatial information while maintaining low computational complexity.

[0102] The structure of GASA is as follows Figure 5 As shown, the input feature map of GASA is x∈R C×H×W First, the input is normalized and channel-expanded. The expanded feature map is divided into two parts along the channel dimension. The first C channel features r1 and r2 are respectively subjected to depthwise separable convolution along the vertical and horizontal axes to obtain response features in two directions. The two axial outputs are summed at each spatial position (i, j) to form an attention map. It is expressed as follows:

[0103]

[0104] where r ij ∈R d is the attention vector at position (i, j), and Represents the axial convolution kernel fused with relative position encoding, corresponding to the width and height responses respectively. The last C channel features X* are used as feature branches, which are modulated by the above attention map and then gated element by element. The results are fused through 1×1 convolution and a learnable scaling parameter γ∈R is introduced. 1×C×1×1 , and finally added to the original input residual to get the final output:

[0105] y ij =γ·ψ(x ij ⊙sigmoid(a ij ))+x ij (4)

[0106] where x ij represents the original input features, and ψ(·) is the channel-fused convolution function. This module structure can extract spatial semantic information through the convolution receptive field without explicitly constructing query-key interaction, and enhance it with a gating mechanism.

[0107] Technical effects:

[0108] The differences in canopy structure and branch density are captured through vertical and horizontal convolution respectively, thereby improving the discrimination of tree age-related features.

[0109] Dynamically modulate the characteristic response intensity to effectively reduce interference from background areas such as bare soil and roads.

[0110] Preserve the spatial details of the original input and avoid the vanishing gradient problem in deep network training.

[0111] In another technical solution, the gated axial spatial pyramid structure constructs a multi-scale spatial modeling mechanism for Wogan fruit tree plots by layering the gated axial spatial attention modules, specifically including: processing the input features through multiple parallel branches, each branch compressing the channel dimension through 1×1 convolution and then introducing 1-4 layers of the gated axial spatial attention modules, and gradually expanding the receptive field using modules of different layers, wherein the shallow branches focus on the fine-grained canopy texture of the Wogan seedling plots through single-layer modules, and the deep branches capture the wide-area canopy distribution characteristics of the 4-5 year old plots through multi-layer modules; the multi-scale features output by the module are unified in resolution through bilinear interpolation, and then spliced ​​with the shallow detail features extracted by the MobileNetV2 backbone network, fusing the local branch density information and the global spatial layout characteristics of the Wogan fruit tree plots, thereby solving the boundary fuzzy problem caused by morphological differences of plots of different ages under the multi-age mixed planting mode, and suppressing the misclassification interference of young plots and bare soil areas.

[0112] The input feature map is processed by four parallel branches, and each branch compresses the number of channels to 1 / 4 of the original channels through 1×1 convolution. For example, when the input feature map has 256 channels, each branch is compressed to 64 channels. The first branch introduces a 1-layer gated axial spatial attention (GASA) module, the second branch introduces 2 layers, the third branch introduces 3 layers, and the fourth branch introduces 4 layers. The axial convolution kernel size of each layer of GASA module can be set to 3×1 (vertical axis) and 1×3 (horizontal axis) to gradually expand the receptive field. The shallow branches (1-2 layers) focus on the fine-grained texture of the seedling plots, and the deep branches (3-4 layers) capture the wide-area distribution of the adult plots.

[0113] The parallel branch can be implemented based on PyTorch's nn.ModuleList, and the 1×1 convolution layer can use the standard convolution module.

[0114] Workflow: The input feature map enters four branches, where it is channel-compressed and stacked with GASA modules layer by layer. The spatial attention mechanism of GASA modules at different levels dynamically expands the receptive field, capturing long-range dependencies and deep details. The spatial resolution of the output feature maps of each branch remains consistent.

[0115] The feature maps output by the gated axial spatial pyramid structure are bilinearly interpolated to the same resolution as the shallow features of MobileNetV2. For example, when the deep branch outputs a 128×128 resolution, they are interpolated and upsampled to 512×512. Shallow features are extracted from the third inverted residual block of MobileNetV2, retaining a resolution of 1 / 4 the original input size. During fusion, multi-scale features are concatenated along the channel dimension. For example, the four branches each output 64 channels, which are concatenated to 256 channels. This is then concatenated with the 128 channels of the shallow features to form 384 channels.

[0116] Bilinear interpolation can be performed by calling PyTorch's F.interpolate function with the upsampling mode set to bilinear. Feature concatenation is implemented using torch.cat, and memory management relies on CUDA memory optimization.

[0117] Workflow: After interpolation and alignment, the feature maps of each branch are concatenated with shallow features extracted by MobileNetV2 (such as branch density and soil texture). The fused features are input into two 3×3 convolutional layers. The first layer outputs 256 channels, and the second layer outputs 128 channels. The activation function is ReLU.

[0118] The cross-scale fused features are restored to the original image resolution by four times upsampling, and the upsampling method is bilinear interpolation. The classification head uses 1×1 convolution to compress the number of channels to the number of categories (such as 5 categories), outputs the category score map of each pixel, and finally uses the Softmax activation function to normalize each pixel in the category dimension to obtain the probability distribution of each category for generating the segmentation result.

[0119] The classification head convolution layer can be implemented based on PyTorch's nn.Conv2d, and the Softmax function is normalized along the channel dimension.

[0120] Workflow: After upsampling the fused features, the classification head generates per-pixel class probabilities. The probability map undergoes threshold segmentation and morphological optimization, outputting a segmentation result with continuous boundaries that is directly mapped to the drone orthophoto coordinate system.

[0121] Technical effects:

[0122] 1. Multi-scale feature capture: By layering and stacking GASA modules, we adapt to the morphological gradient characteristics of plots from seedling to adult stage, and enhance the joint modeling capabilities of fine texture and wide-area features.

[0123] 2. Resolution alignment and fusion: Bilinear interpolation eliminates scale offset, complements shallow details with deep semantics, and enhances boundary segmentation consistency in mixed planting scenarios.

[0124] 3. Applicability of classification results: Pixel-level matching output can be directly imported into orchard management systems, supporting accurate decision-making for differentiated fertilization and pest and disease control.

[0125] In another technical solution, the fusion mechanism of the multi-scale features output by the gated axial spatial pyramid structure and the shallow features is designed for the morphological complexity of the Wogan fruit tree plots, specifically including the following steps:

[0126] Step 1: Shallow feature extraction and retention: Based on the shallow features extracted by the MobileNetV2 backbone network, the fine-grained branch density distribution characteristics and canopy edge texture details of the Wogan seedling plots are retained to enhance the ability to distinguish between young fruit tree plots and bare soil and road non-plot areas;

[0127] Step 2: Multi-scale Feature Alignment and Adaptation: A bilinear interpolation algorithm is used to upsample the multi-scale features (shallow branches: fine textures in seedlings; deep branches: wide-area canopy layers in 4-5-year-old trees) output by the gated axial spatial pyramid to the same spatial resolution as the shallow features extracted by the MobileNetV2 backbone network (e.g., 128×128), thereby achieving geometric alignment of cross-scale features. This pixel-by-pixel alignment eliminates the effect of feature scale shift caused by differences in tree age gradients. This ensures that fine-grained features such as local branch density distribution and canopy edge texture in seedling plots are geometrically consistent with global spatial features such as wide-area canopy layout and intersecting boundaries of adjacent plots in mature plots before semantic fusion. This mechanism effectively suppresses boundary blurring caused by feature resolution mismatch in multi-age mixed planting scenarios, providing a precise spatial reference for cross-scale feature fusion, thereby improving the model's ability to model tree age gradients and the geometric fidelity of classification boundaries.

[0128] Step 3: Cross-scale feature fusion and boundary reinforcement: The aligned multi-scale features and shallow features are spliced ​​along the channel dimension, and cross-scale information fusion is performed through two 3×3 convolutional layers:

[0129] First convolution: Extract the branch density gradient variation characteristics within the Wogan fruit tree plots and enhance the differential response of canopy cover between plots of different tree ages;

[0130] Secondary convolution: Captures the patchy distribution characteristics of plots sprayed with lime and pesticides and the scattered boundary mutation signals of young fruit tree plots, and suppresses the missegmentation of adjacent plots in mixed planting patterns due to canopy overlap or background interference;

[0131] Step 4: Reconstructing a high-resolution classification map: The fused features are restored to the original image resolution through four-fold upsampling to generate a classification result map that matches the pixel-level of the drone orthophoto. This accurately marks the continuous boundaries and internal structure of the mandarin orange tree plots of different ages, solving the problems of blurred boundaries caused by mixed planting of multiple ages and missed detection of transition areas between young plots and bare soil.

[0132] Shallow features of the MobileNetV2 backbone network are extracted from the third inverted residual block, with an output resolution of 1 / 4 the input image. For example, when inputting a 512×512 image, the shallow feature size is 128×128×64. These features include fine-grained branch density distribution (≥15 branches per square meter) and canopy edge texture (edge ​​gradient ≥0.2) for seedling plots.

[0133] Feature extraction can be implemented based on the PyTorch framework, and the pre-trained weights of MobileNetV2 can be loaded from the public model library.

[0134] The model's multi-scale features include shallow features extracted from the early and middle stages of the main network (resolution 128×128, 64 channels), focusing on capturing fine-grained texture information in seedling-stage plots; and deep features derived from the ASPP module (resolution 64×64, 128 channels), which characterize the broad distribution of mature vegetation. To achieve spatial alignment, the deep features are upsampled to 128×128 using PyTorch's bilinear interpolation mode, with an interpolation error of ±1 pixel. Before concatenation, L2 normalization (ε=1e-6) is applied to each branch along the channel dimension to eliminate amplitude shifts caused by scale differences and improve multi-scale fusion.

[0135] Bilinear interpolation can be integrated into the model's forward propagation process. Normalization layers are implemented using PyTorch's nn.BatchNorm2d. Feature concatenation uses the torch.cat function, and memory usage is optimized using gradient checkpointing.

[0136] Workflow: After the deep branch features are interpolated and amplified, they are concatenated with the shallow branch features along the channel dimension (64 + 128 = 192 channels). Normalization is applied independently to each branch before concatenation to ensure feature distribution consistency.

[0137] The first 3×3 convolutional layer outputs 128 channels, with kernel weights initialized to a He normal distribution and biases initialized to 0. The activation function used is LeakyReLU (negative slope = 0.01) to preserve weak negative responses. The second 3×3 convolutional layer outputs 64 channels to capture the discontinuous boundaries of pesticide-sprayed patches (minimum patch area ≥ 50 pixels). For classification map reconstruction, quadruple upsampling is achieved via two bilinear interpolations (2x×2x). After interpolation, a 1×1 convolution is performed to adjust the channels to the number of categories. The softmax temperature parameter is set to 1.0.

[0138] In another technical solution, the construction of the GACL-DeepLabV3+ model is designed for a multi-age mixed planting pattern of Wogan fruit trees, specifically including the following steps:

[0139] Step 1: Multi-level feature extraction of input images: The visible light image of the Wogan fruit tree plot acquired by the drone is input into the MobileNetV2 backbone network to extract shallow features and high-level features respectively:

[0140] Shallow layer characteristics: retain the fine-grained shoot density distribution, canopy edge texture of seedling-stage plots, and soil characteristics of exposed surfaces of young plots to distinguish non-fruit tree areas;

[0141] High-level features: Capture the wide-area canopy distribution pattern of 4-5 year old plots and the patchy spectral response of pesticide-sprayed plots, characterizing the global spatial layout of mature Wogan fruit trees;

[0142] Step 2: Channel-aware attention guides spatial feature enhancement: The channel-aware lightweight attention module is embedded after the high-level features to optimize the semantic distinction ability of the Wogan plot classification through the following operations:

[0143] Extract channel semantic information of high-level features based on global average pooling and generate channel semantic guided key vectors;

[0144] Convert the spatial location information into a query matrix and calculate the correlation weight between channel semantics and spatial location;

[0145] Step 3: Multi-scale spatial modeling and tree age feature adaptation: The high-level features enhanced by CLA are input into the gated axial spatial pyramid structure to adapt to the age differences of the Wogan fruit tree plots through the following mechanism:

[0146] Shallow branching focused on the fine-grained canopy texture of the seedling-stage plots, enhancing the contrast between young fruit trees and the bare soil boundary;

[0147] Deep branches expand the receptive field, capturing the wide-area canopy distribution of 4-5 year old tree plots and the staggered boundary characteristics between adjacent plots;

[0148] The multi-scale features output by each branch are aligned by bilinear interpolation to retain the morphological gradient characteristics of plots of different tree ages;

[0149] Step 4: Classification result generation and application optimization: The cross-scale fused features are reconstructed through a high-resolution method to generate a classification result map that matches the pixel level of the drone orthophoto. The continuous boundaries and transition areas of the multi-age mandarin orange tree plots are accurately marked, and the results are imported into the orchard intelligent management system to guide differentiated fertilization, pest and disease control, and harvest planning, significantly reducing manual survey costs and improving orchard management efficiency.

[0150] Shallow features of the MobileNetV2 backbone network are extracted from the third inverted residual block, with an output resolution of 1 / 4 the input image (e.g., 512×512 input corresponds to 128×128 output) and 64 channels. Shallow features include the branch density distribution (branch spacing ≤ 15 cm) and canopy edge gradients (Sobel operator gradient amplitude ≥ 0.3) for seedling-stage plots. High-level features are extracted from the last layer of MobileNetV2, with a resolution of 32×32 and 320 channels, capturing the wide-area canopy distribution (canopy diameter ≥ 2 m) for mature-stage plots.

[0151] Feature extraction can be implemented based on the PyTorch framework, and the pre-trained weights of MobileNetV2 can be loaded from the open source library.

[0152] Workflow: High-level features are input into the CLA module to generate a spatial attention map, which is weighted and then connected to the original feature residual.

[0153] Multi-scale features are unified to a 128×128 resolution via bilinear interpolation and concatenated with shallow features to create 384 channels (64+320). The fused features are fed into two 3×3 convolutional layers, with the first layer outputting 256 channels and the second layer outputting 128 channels. The classification result map is then restored to a 512×512 resolution via quadruple upsampling and exported as a GeoTIFF, with the coordinate system consistent with the drone orthophoto. The orchard management system interface supports GeoJSON format, and classification results are uploaded via a REST API to guide variable operation of the fertilizer applicator (fertilizer application gradient: 20 kg / mu in the seedling stage, 50 kg / mu in the adult stage).

[0154] Workflow: The fused features are convolutionally compressed and then upsampled to generate a classification map, which is then converted into a geospatial data format. The results are uploaded to the orchard management system, which automatically generates a differentiated fertilization prescription map to drive the agricultural machinery.

[0155] Technical effects:

[0156] 1. Multi-level feature complementarity: Jointly modeling shallow details and high-level semantics improves the ability to distinguish between seedling and mature plots.

[0157] 2. Attention-guided optimization: Co-weighting of channel semantics and spatial position to suppress background interference and enhance pesticide patch response.

[0158] 3. Application system compatibility: Geospatial data formats are seamlessly integrated with agricultural machinery systems, enabling closed-loop management from classification results to field operations.

[0159] <Application Examples>

[0160] The study area (geographic coordinates: 22°59′58″N~23°33′16″N, 107°49′26″E~108°37′22″E) in the Wogan mandarin orange planting area of ​​Luwo Town, Wuming District, Nanning City, Guangxi Zhuang Autonomous Region, China was selected as the research area. Using the DJI Mavic 3 multispectral drone and the GACL-DeepLabV3+ model, we achieved fine classification of multi-age Wogan mandarin orange tree plots. The specific steps are as follows:

[0161] Step 1: UAV image data collection

[0162] Data collection was performed using a DJI Mavic 3 multispectral drone, with a pre-set flight path ensuring complete image coverage of the study area. Flight parameters were set as follows: an altitude of 250 meters to balance image resolution (0.0685 m) with coverage efficiency; 80% heading overlap and 70% lateral overlap to ensure image stitching accuracy; data collection was performed between 10:00 AM and 2:00 PM on sunny days, when light intensity ranged from 50,000 to 80,000 lux to reduce shadow interference; and visible light imagery encompassing red, green, and blue wavelengths was collected to enhance the differentiation of vegetation features.

[0163] Through image sensors and optimized flight parameters, the original images obtained have high resolution and low noise characteristics, providing a high-quality data foundation for subsequent pre-processing.

[0164] Step 2: Data preprocessing and dataset construction

[0165] 2.1 Image Preprocessing

[0166] The original images were processed using Pix4Dmapper (v4.5.6) software. First, feature extraction and aerial triangulation were performed to generate a dense point cloud to achieve geometric alignment between images. A digital surface model (DSM) was constructed, and orthorectification was performed to generate a high-precision orthophoto (DOM) with a resolution of 0.0685 m.

[0167] 2.2 Manual Labeling and Dataset Division

[0168] Under the guidance of orchard experts and combined with field investigations, the Wogan fruit tree plots were divided into five categories: seedling plots, one-year-old plots, 2-3-year-old plots, 4-5-year-old plots, and plots sprayed with lime and pesticides; the Labelme tool was used to manually annotate the orthophotos with polygons; the annotated images were cut into 512×512 pixel sizes, and Gaussian noise injection (intensity 15, 25, 40) and rotation transformation (90° / 180° / 270°) were used to simulate real-world interference, ultimately generating 2032 slice images; the dataset was divided into a training set (1626 images), a validation set (203 images), and a test set (203 images) in an 8:1:1 ratio to ensure balanced category distribution.

[0169] The preprocessing process improves image quality through geometric correction and noise suppression, while the labeling and enhancement strategies construct a multi-scale, interference-resistant dataset based on the characteristics of Wogan cultivation scenarios, laying the foundation for model training.

[0170] Step 3: Channel-aware lightweight attention module (CLA) design

[0171] We propose a channel-aware lightweight attention (CLA), the core of which is to guide the generation of spatial attention weights through channel semantics to achieve response enhancement of spatially salient regions. The specific process is as follows:

[0172] Channel semantic extraction: Perform global average pooling on the input feature map x along the spatial dimension to obtain the overall response vector Z of each channel. The channel semantic-guided Key vector K is obtained through linear transformation. The calculation formula is:

[0173]

[0174] This vector is transformed by the linear transformation w k ∈R C'×C Get the channel semantics guided Key, denoted as K∈W k ·K c ∈R C′ .

[0175] Spatial Query Generation: Flatten the input feature x into a matrix of H×W×C, and obtain the Query matrix Q through linear transformation, that is, the input feature x is flattened into x * ∈R (H·W)×C , and through the linear transformation w q ∈R C'×C Get the Query matrix Q∈W q ·x * R C ′×(H·W) .

[0176] Attention map calculation: Generate the spatial attention map A by normalizing the dot product of Key and Query. The formula is:

[0177] A=soft max(K T ·Q)∈R 1×(H·W ) (1)

[0178] Feature weighting and residual connection: Multiply the attention map with the original feature element by element, and output the enhanced feature through the residual connection. The formula is: y = γ·A⊙x+x (2)

[0179] Here, γ∈R is a learnable scaling factor.

[0180] Technical function: CLA models the correlation between channel semantics and spatial location, enhancing the model's responsiveness to key areas and reducing feature confusion among similar plots.

[0181] Step 4: Gated Axial Spatial Attention Module (GASA) Design

[0182] In order to enhance the model's ability to model spatial structure and target category boundaries in semantic segmentation tasks, a lightweight gated axial spatial attention (GASA) is proposed. The core of GASA is to combine the idea of ​​axial attention on the basis of spatial attention and introduce learnable gating parameters to enhance the perception of local and long-range spatial information while maintaining low computational complexity.

[0183] The input feature map of GASA is x∈R C×H×W First, the input is normalized and channel-expanded. The expanded feature map is divided into two parts along the channel dimension. The first C channel features r1 and r2 are subjected to depthwise separable convolution along the vertical and horizontal axes respectively to obtain response features in two directions. The two axial outputs are summed at each spatial position (i, j) to form an attention map aij, which is expressed as follows:

[0184]

[0185] where r ij ∈R d is the attention vector at position (i, j), and Represents the axial convolution kernel fused with relative position encoding, corresponding to the responses of width and height respectively.

[0186] The last C channel features X* are used as feature branches, which are modulated by the above attention map and then gated element by element. The results are fused through 1×1 convolution and a learnable scaling parameter γ∈R is introduced. 1×C×1×1 , and finally added to the original input residual to get the final output:

[0187] y ij =γ·ψ(x ij ☉sigmoid(a ij ))+x ij ;

[0188] where x ij represents the original input features, ψ(·) is the convolution function of channel fusion, and this module structure does not require explicit construction of query-key interaction.

[0189] Technical function: GASA separates spatial modeling and feature modulation paths, reducing computational complexity (FLOPs reduced by 30%) while improving the continuity of boundary area features. It is particularly suitable for contour segmentation of old tree plots (such as 4-5 years old trees).

[0190] Step 5: Gated Axial Space Pyramid (GASP) design

[0191] GASP replaces the traditional ASPP's dilated convolution with a GASA module to construct a multi-scale spatial modeling path. Its design includes: Four parallel branches, each of which compresses the channel via 1×1 convolutions and then introduces 1-4 layers of GASA modules. By increasing the number of GASA layers, a multi-level receptive field from local to global is constructed, overcoming the ASPP's fixed expansion rate limitations. The outputs of each branch are uniformly resolved through bilinear interpolation, then concatenated with the shallow features of the backbone network along the channel axis and fused through two 3×3 convolutions.

[0192] Technical advantages: GASP adaptively captures differences in land morphology through a dynamic attention mechanism, reduces misjudgments caused by background interference, and significantly reduces edge segmentation errors compared to ASPP.

[0193] Step 6: GACL-DeepLabV3+ model construction

[0194] The model is based on the DeepLabV3+ framework and integrates a lightweight backbone network and attention module: MobileNetV2 is used to replace Xception, and the parameter size is compressed to 5.876MB through inverse residual blocks and linear bottleneck design, which is suitable for drone deployment; the CLA module is inserted after the high-level features of MobileNetV2 to guide the model to focus on key areas; the features processed by CLA are input into the GASP structure to generate multi-scale spatial perception features; finally, the GASP output is spliced ​​with the shallow features, and the classification result image is output through two 3×3 convolutions and four times upsampling.

[0195] Model advantages: The lightweight design combined with the attention mechanism improves classification accuracy (mIoU 94.03%) while ensuring real-time performance.

[0196] Step 7: Model training and classification verification

[0197] 7.1 Training Configuration

[0198] The hardware environment is Linux system, NVIDIA A100 GPU and PyTorch framework; the optimization strategy adopts stochastic gradient descent (SGD) optimizer, with an initial learning rate of 7e-3 and a learning rate decrease method of cosine annealing; the loss function is cross entropy loss, which balances classification accuracy and boundary continuity; the training parameters are set to Batchsize = 4 and Epoch = 100.

[0199] 7.2 Classification Verification

[0200] Input the test set into the trained model and output the classification results as follows Figure 8-12 The performance metrics are: mean Intersection over Union (mIoU) of 94.03%, mean pixel accuracy (mPA) of 97.02%, mean precision (mPrecision) of 96.82%, and overall accuracy (OA) of 96.71%. The results show that the model achieves high classification accuracy for seedling plots (low vegetation cover) and plots sprayed with lime and pesticides (spectral anomalies), effectively addressing the challenges of blurred boundaries and similar features encountered in traditional methods.

[0201] The present invention proposes a method for fine classification of mandarin orange tree plots based on UAV remote sensing images using GACL-DeepLabV3+, which can effectively solve the classification and extraction difficulties caused by the blurred boundaries of mandarin orange tree plots of different ages and states in UAV remote sensing images, the similarity of characteristics with other tree species and crops, and the overlapping characteristics of exposed surfaces and non-plot areas (such as dirt roads, wasteland, etc.) of young tree-planted plots, thereby realizing the fine classification of mandarin orange tree plots of different ages and states.

[0202] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.

Claims

1. A fine classification method for mandarin orange tree plots based on UAV remote sensing images using GACL-DeepLabV3+, characterized by: The following steps are involved: S1: Use UAV to obtain visible light remote sensing images of mandarin orange fruit trees of different ages and conditions in the study area; S2: Preprocessing the acquired remote sensing images to generate centimeter-level orthophotos. The preprocessing includes feature extraction and aerial triangulation to generate dense point clouds, digital surface model construction, and orthorectification. S3: Construct a dataset of mandarin orange tree plots based on orthophotos. The dataset includes plots in the seedling stage, plots with 1-year-old fruit trees, plots with 2-3-year-old fruit trees, plots with 4-5-year-old fruit trees, and plots with lime-sprayed fruit trees. The images are annotated using manual annotation tools and then segmented. S4: Perform data augmentation on the dataset, including applying Gaussian noise with an intensity of 10-50 and rotation transformation to the image to generate a data augmented training set with noise robustness and multi-angle features, and divide it into training set, validation set and test set in proportion; S5: Construct a GACL-DeepLabV3+ model, which includes a MobileNetV2 backbone network, a channel-aware lightweight attention module, a gated axial spatial attention module, and a gated axial spatial pyramid structure. The GACL-DeepLabV3+ model is trained using a data augmentation training set. A stochastic gradient descent optimizer is used during training, with an initial learning rate of 7e-3, a cosine annealing learning rate decay method, a batch size of 4, and 100 training rounds. S6: Input the test set into the trained GACL-DeepLabV3+ model and output the classification results of the Wogan fruit tree plots.

2. The method for fine classification of mandarin orange tree plots based on UAV remote sensing images using GACL-DeepLabV3+ according to claim 1, wherein The operation of the drone to obtain visible light remote sensing images further includes: establishing a flight parameter configuration table based on a preset fruit tree age gradient database and canopy morphological feature library, using a multi-rotor drone equipped with a visible light multispectral imager with a polarization filtering function, and performing aerial photography with the canopy as the reference plane during the morning scattered light-dominated period. Young trees with a tree age of less than 3 years are photographed at a flight altitude of 250 meters and an 80% heading overlap rate, while mature and high-yield trees are photographed at a flight altitude of 250 meters and a 70% heading overlap rate. An edge computing unit-based real-time image quality diagnosis module is configured to perform online detection and re-photography decisions on leaf texture clarity and abnormal fruit coloring areas.

3. The method for fine classification of mandarin orange tree plots based on UAV remote sensing images using GACL-DeepLabV3+ according to claim 1, wherein The manual labeling tool is Labelme, and the labeled image segmentation adopts an adaptive segmentation algorithm based on saliency guidance.

4. The method for fine classification of mandarin orange tree plots based on UAV remote sensing images using GACL-DeepLabV3+ according to claim 1, wherein The channel-aware lightweight attention module extracts the channel semantic information of the input features through global average pooling, generates a channel semantics-guided Key vector, and generates a spatial Query matrix through linear transformation. The Key vector and the Query matrix are normalized by dot product to generate a spatial attention map, which is then weighted and output through residual connections.

5. The method for fine classification of mandarin orange tree plots based on UAV remote sensing images using GACL-DeepLabV3+ according to claim 1, wherein: The gated axial spatial attention module enhances the spatial feature modeling capability of the Wogan fruit tree plots through axial convolution and gating mechanisms, specifically including: dividing the input features into two parts along the channel dimension, applying depthwise separable convolution along the longitudinal and transverse axes to the first C channel features, capturing the canopy structure and branch density differences of the Wogan fruit tree plots in the vertical and horizontal directions, generating axial response features and summing them, and adaptively modulating the response features of the last C channel features through learnable gating parameters to suppress background interference and enhance the boundary features of plots of different tree ages, and then outputting the fused features through residual connections, thereby solving the problem of blurred boundaries of Wogan fruit tree plots caused by mixed planting and overlapping features of young plots and non-plot areas.

6. The method for fine classification of mandarin orange tree plots based on UAV remote sensing images using GACL-DeepLabV3+ according to claim 1, wherein: The gated axial spatial pyramid structure constructs a multi-scale spatial modeling mechanism for Wogan fruit tree plots by stacking the gated axial spatial attention modules in layers, specifically including: processing the input features through multiple parallel branches, each branch compressing the channel dimension through 1×1 convolution and then introducing 1-4 layers of the gated axial spatial attention modules, and gradually expanding the receptive field using modules of different layers, wherein the shallow branches focus on the fine-grained canopy texture of the Wogan seedling plots through single-layer modules, and the deep branches capture the wide-area canopy distribution characteristics of the 4-5 year old plots through multi-layer modules; the output multi-scale features are unified in resolution through bilinear interpolation, and then spliced ​​with the shallow detail features extracted by the MobileNetV2 backbone network, fusing the local branch density information and the global spatial layout characteristics of the Wogan fruit tree plots, thereby solving the boundary fuzzy problem caused by morphological differences of plots of different ages in the multi-age mixed planting mode, and suppressing the misclassification interference of young plots and bare soil areas.

7. The method for fine classification of mandarin orange tree plots based on UAV remote sensing images using GACL-DeepLabV3+ according to claim 6, characterized in that: The fusion mechanism of the multi-scale features output by the gated axial spatial pyramid structure and the shallow features is designed based on the morphological complexity of the Wogan fruit tree plots, and specifically includes the following steps: Step 1: Shallow feature extraction and retention: Based on the shallow features extracted by the MobileNetV2 backbone network, the fine-grained branch density distribution characteristics and canopy edge texture details of the Wogan seedling plots are retained to enhance the ability to distinguish between young fruit tree plots and bare soil and road non-plot areas; Step 2: Multi-scale feature alignment and adaptation: A bilinear interpolation algorithm is used to upsample the multi-scale features output by the gated axial spatial pyramid to ensure that their spatial resolution is consistent with the shallow features extracted by the MobileNetV2 backbone network, thereby achieving geometric alignment of cross-scale features. Step 3: Cross-scale feature fusion and boundary reinforcement: The aligned multi-scale features and shallow features are spliced ​​along the channel dimension, and cross-scale information fusion is performed through two 3×3 convolutional layers: First convolution: Extract the branch density gradient variation characteristics within the Wogan fruit tree plots and enhance the differential response of canopy cover between plots of different tree ages; Secondary convolution: Captures the patchy distribution characteristics of plots sprayed with lime and pesticides and the scattered boundary mutation signals of young fruit tree plots, and suppresses the missegmentation of adjacent plots in mixed planting patterns due to canopy overlap or background interference; Step 4: Reconstructing a high-resolution classification map: The fused features are restored to the original image resolution through four-fold upsampling to generate a classification result map that matches the pixel-level of the drone orthophoto. This accurately marks the continuous boundaries and internal structure of the mandarin orange tree plots of different ages, solving the problems of blurred boundaries caused by mixed planting of multiple ages and missed detection of transition areas between young plots and bare soil.

8. The method for fine classification of mandarin orange tree plots based on UAV remote sensing images using GACL-DeepLabV3+ according to claim 7, characterized in that: The construction of the GACL-DeepLabV3+ model is designed for a multi-age mixed planting pattern of Wogan fruit trees, specifically including the following steps: Step 1: Multi-level feature extraction of input images: The visible light image of the Wogan fruit tree plot acquired by the drone is input into the MobileNetV2 backbone network to extract shallow features and high-level features respectively: Shallow layer characteristics: retain the fine-grained shoot density distribution, canopy edge texture of seedling-stage plots, and soil characteristics of exposed surfaces of young plots to distinguish non-fruit tree areas; High-level features: Capture the wide-area canopy distribution pattern of 4-5 year old plots and the patchy spectral response of pesticide-sprayed plots, characterizing the global spatial layout of mature Wogan fruit trees; Step 2: Channel-aware attention guides spatial feature enhancement: The channel-aware lightweight attention module is embedded after the high-level features to optimize the semantic distinction ability of the Wogan plot classification through the following operations: Extract channel semantic information of high-level features based on global average pooling and generate channel semantic guided key vectors; Convert the spatial location information into a query matrix and calculate the correlation weight between channel semantics and spatial location; Residual connections are used to enhance the response intensity to sparse shoot areas in seedling plots and local patches of pesticide spraying plots, thereby suppressing background noise interference. Step 3: Multi-scale spatial modeling and tree age feature adaptation: The high-level features enhanced by CLA are input into the gated axial spatial pyramid structure to adapt to the age differences of the Wogan fruit tree plots through the following mechanism: The multi-layer branching structure expands the receptive field, deeply extracts data texture features, enhances the contrast between young fruit trees and bare soil boundaries, and captures the wide-area canopy distribution of 4-5 year old tree plots and the staggered boundary features between adjacent plots. The output multi-scale features are aligned by bilinear interpolation to preserve the morphological gradient characteristics of plots of different tree ages; Step 4: Classification result generation and application optimization: The cross-scale fused features are reconstructed through a high-resolution method to generate a classification result map that matches the pixel level of the drone orthophoto. The continuous boundaries and transition areas of the multi-age mandarin orange tree plots are accurately marked, and the results are imported into the orchard intelligent management system to guide differentiated fertilization, pest and disease control, and harvest planning, significantly reducing manual survey costs and improving orchard management efficiency.

Citation Information

Patent Citations

  • Pear tree planting area remote sensing extraction method based on Re-UNet model

    CN117636170A

Cited By

  • Method for detecting control effect of mikania micrantha based on aerial photography of unmanned aerial vehicle

    CN122049740A

  • A method for detecting the control effect of Mikania micrantha based on drone aerial photography

    CN122049740B